Submitted:
28 July 2026
Posted:
03 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We establish formal definitions of agent system self-improvement and RSI, together with a five-level capability grading standard ranging from manual improvement to general RSI.
- We propose a unified research framework (Figure 1) that models the foundation model, agent harness, agent data system, agent trainer, and improvement mechanism as a coupled evolving system, capturing the scope, dependencies, and dynamics of self-improvement.
- We systematically reviews existing work on self-improving agent systems and provides a taxonomy (Figure 2), covering self-improvement in individual components and co-improvement across components.
- We identify key open problems on the path to RSI, which provide research directions for developing more autonomous, reliable, and general self-improving agent systems.
2. Foundation of Self-Improving Agent Systems
2.1. Preliminaries
2.1.1. Agent Systems
Foundation model.
Agent harness.
Agent data system.
Agent trainer.
Improvement mechanism.
2.1.2. Self-Improvement
Recursive Self-Improvement.
2.2. Self-Improvement Capability Grading Standard
- (L1)
- Manual improvement. The agent system has no autonomous improvement capability. It is deployed and executed in a fixed form, and every change requires an external development and deployment process.
- (L2)
- Assisted improvement. The agent system can propose candidate modifications or provide diagnostic evidence, but humans remain responsible for validating and applying substantive changes. The improvement mechanism is maintained by humans rather than by the agent system. Therefore, L2 is regarded as a precursor to self-improvement rather than an instance of it.
- (L3)
- Programmatic self-improvement. The agent system can autonomously propose, apply, and validate modifications to its operational components, such as the foundation model, agent harness, data system, or trainer. These modifications may be substantial and may span multiple components. However, the improvement mechanism that governs these modifications remains fixed or externally maintained. L3 is therefore programmatic: the system controls the execution of improvement, but does not yet modify the mechanism that governs subsequent improvement cycles.
- (L4)
- Bounded recursive self-improvement. The agent system can not only propose candidate modifications, validate their effects, and apply the validated modifications, but also rewrite its own improvement mechanism, which makes the improvement mechanism self-referential. Individual steps need not modify every component, but the overall process supports sustained long-term progress within a bounded domain.
- (L5)
- General recursive self-improvement. L5 retains the autonomy, self-reference, and long-term progress of L4. Beyond this, its improvement capability can transfer effectively across broad and evolving task domains rather than remaining limited to a fixed benchmark, narrow task family, or restricted operational setting. Such generality marks a qualitative shift: the system’s capacity to improve is no longer tied to any particular problem space, opening the possibility of improvement across an ever-expanding range of domains.
2.3. Unified Research Framework for Agent System Self-Improvement
2.3.1. Definition and Formulation
System state.
Operational dependencies.
Self-Improvement Process.
- 1.
- The Diagnose stage analyzes end-to-end performance, component states, interface behavior, and feedback signals to identify the bottlenecks that currently constrain system-level improvement.
- 2.
- The Propose stage generates candidate modifications, each specifying the target subset of and the proposed change. A candidate may target one component or coordinate changes across multiple interdependent components.
- 3.
- The Evaluate stage assesses the effects of candidate modifications on the coupled system, including performance, cost, compatibility, and safety, and selects candidates for integration.
- 4.
- The Integrate stage applies accepted modifications to produce while preserving cross-component consistency and ensuring that the changes persist across subsequent improvement cycles.
2.3.2. Advantages of the Unified Framework
3. Agent Harness Self-Improvement
| Work | Self-Improvement Process | Level | ||
| Object | Evidence | Mechanism | ||
| Module-Level Self-Improvement | ||||
| Reflexion [26] | Experience memory | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| SWE-Exp [28] | Experience memory | Trace + Outcome label | Experience reuse | L3 |
| ReasoningBank [29] | Experience memory | Trace + Evaluator judgment | Experience reuse | L3 |
| FLEX [30] | Experience memory | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| Decocted Experience [31] | Experience memory | Trace + Outcome label | Experience reuse | L3 |
| Live-Evo [41] | Experience memory | Trace + Outcome label | Experience reuse | L3 |
| Memory Transfer Learning [55] | Experience memory | Cross-domain feedback | Experience reuse | L3 |
| Tool Makers [57] | Tool library | Execution result + Outcome label | Artifact synthesis | L3 |
| ToolCoder [58] | Tool library | Execution result | Artifact synthesis | L3 |
| SkillWeaver [60] | Skill library | Trace + Execution result + Evaluator judgment | Artifact synthesis | L3 |
| SkillX [61] | Skill library | Trace + Outcome label + Evaluator judgment | Artifact synthesis | L3 |
| Workflow-to-Skill [62] | Skill library | Trace | Artifact synthesis | L3 |
| SkillFoundry [66] | Skill library | Execution result + Outcome label + Evaluator judgment | Artifact synthesis | L3 |
| OpenSkill [65] | Skill library | Execution result + Outcome label + Evaluator judgment | Artifact synthesis | L3 |
| SkillOS [69] | Skill library | Trace + Evaluator judgment | Policy learning | L3 |
| SkillComposer [74] | Skill library | Trace + Outcome label | Policy learning | L3 |
| Programmatic Skill Networks [75] | Skill library | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| SkillAxe [76] | Skill library | Evaluator judgment | Diagnostic repair | L3 |
| SkillAudit [77] | Skill library | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| SkillGen [78] | Skill library | Trace + Outcome label + Evaluator judgment | Artifact synthesis | L3 |
| CoEvoSkills [79] | Skill library | Execution result + Outcome label + Evaluator judgment | Artifact synthesis | L3 |
| SkillSmith [83] | Skill-tool library | Trace + Execution result + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| Confucius Code Agent [85] | Skill-tool library | Trace + Execution result + Evaluator judgment | Artifact synthesis | L3 |
| Dynamic Cheatsheet [110] | Context playbook | Evaluator judgment | Experience reuse | L3 |
| ACE [111] | Context playbook | Trace + Execution result + Evaluator judgment | Experience reuse | L3 |
| SCOPE [112] | Context playbook | Trace + Execution result + Evaluator judgment | Experience reuse | L3 |
| Reflective Context Learning [113] | Context playbook | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| Unified Context Evolution [114] | Context playbook | Trace + Outcome label + Improvement statistics | Experience reuse | L3 |
| KACE [115] | Context playbook | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| MEMO [117] | Context playbook | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| SEEK-SQL [118] | Context playbook | Trace + Execution result + Outcome label + Evaluator judgment | Experience reuse | L3 |
| GEPA [103] | Multi-stage prompts | Trace + Execution result + Outcome label | Search-based optimization | L3 |
| Trace2Policy [119] | Decision-rule library | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| Learning to Self-Evolve [121] | Context-editor weights | Trace + Outcome label | Policy learning | L3 |
| SePO [120] | Task prompt + Improver prompt | Outcome label | Search-based optimization + Meta-optimization | L4 |
| Orchestration-Level Self-Improvement | ||||
| AgentGA [129] | Trajectory orchestration | Outcome label + Improvement statistics | Search-based optimization | L3 |
| SWE-Replay [136] | Trajectory memory | Trace + Execution result + Outcome label | Experience reuse | L3 |
| Log-Augmented Generation [140] | Trajectory memory | Trace | Experience reuse | L3 |
| FailureMem [143] | Trajectory memory | Execution result + Outcome label + Evaluator judgment | Experience reuse | L3 |
| EvoRepair [144] | Trajectory memory | Trace + Outcome label + Evaluator judgment | Experience reuse | L3 |
| EvoAgent [152] | Agent composition | Evaluator judgment | Search-based optimization | L3 |
| ADAS [153] | Agent system code | Execution result + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| EvoMAS [158] | Agent composition + Experience memory | Trace + Execution result + Evaluator judgment | Search-based optimization | L3 |
| EVOCHAMBER [159] | Agent composition + Experience memory | Cross-domain feedback | Search-based optimization | L3 |
| MermaidFlow [173] | Workflow structure | Execution result + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| EvoAgentX [174] | Workflow structure | Outcome label | Search-based optimization | L3 |
| EvoFlow [175] | Workflow structure | Outcome label + Evaluator judgment | Search-based optimization | L3 |
| SEW [176] | Workflow structure | Outcome label | Search-based optimization | L3 |
| HyEvo [179] | Workflow structure | Execution result + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| AdaptFlow [180] | Workflow structure | Outcome label + Evaluator judgment | Search-based optimization | L3 |
| JudgeFlow [184] | Workflow structure | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| Lean4Agent / LeanEvolve [232] | Workflow structure | Trace + Execution result + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| EvoFSM [233] | Workflow structure | Trace + Evaluator judgment | Diagnostic repair | L3 |
| ScoreFlow [171] | Workflow-manager weights | Outcome label | Policy learning | L3 |
| Learning to Compose [181] | Workflow-manager weights | Outcome label | Policy learning | L3 |
| Workflow-R1 [183] | Workflow-manager weights | Execution result + Outcome label | Policy learning | L3 |
| Learning to Hand Off [189] | Workflow-manager weights | Outcome label | Policy learning | L3 |
| Self-Referential Code Modification | ||||
| Life-Harness [212] | Harness policy | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| HarnessFix [213] | Harness policy | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| Milkyway [214] | Harness policy | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| Self-Harness [215] | Harness policy | Trace + Execution result + Outcome label | Diagnostic repair | L3 |
| POLARIS [216] | Harness policy | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| DemoEvolve [230] | Runtime scaffold | Trace + Execution result + Outcome label | Search-based optimization | L3 |
| SIGA [231] | Runtime scaffold | Trace + Outcome label | Artifact synthesis | L3 |
| HarnessForge [359] | Runtime scaffold + Model weights | Trace + Outcome label + Evaluator judgment | Diagnostic repair + Policy learning | L3 |
| Agentic Harness Engineering [22] | Runtime scaffold + Evolution history | Trace + Outcome label + Evaluator judgment | Diagnostic repair | L3 |
| Meta-Harness [217] | Runtime scaffold + Evolution history | Trace + Outcome label | Search-based optimization | L3 |
| AgentFlow [218] | Runtime scaffold + Evolution history | Trace + Execution result + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| Adaptive Auto-Harness [239] | Runtime scaffold + Evolution history | Trace + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| Group-Evolving Agents [236] | Agent source code + Evolution history | Trace + Outcome label + Evaluator judgment | Search-based optimization | L3 |
| Gödel Agent [226] | Agent source code + Improvement mechanism | Execution result + Outcome label | Meta-optimization | L4 |
| SICA [227] | Agent source code + Improvement mechanism | Execution result + Outcome label + Evaluator judgment | Search-based optimization + Meta-optimization | L4 |
| Darwin Gödel Machine [234] | Agent source code + Improvement mechanism | Trace + Outcome label + Evaluator judgment | Search-based optimization + Meta-optimization | L4 |
| Huxley-Gödel Machine [235] | Agent source code + Improvement mechanism | Outcome label + Improvement statistics | Search-based optimization | L4 |
| Red Queen Gödel Machine [360] | Agent source code + Evaluation mechanism | Outcome label + Evaluator judgment + Improvement statistics | Search-based optimization + Meta-optimization | L4 |
| HyperAgents [23] | Agent source code + Improvement mechanism | Outcome label + Improvement statistics | Search-based optimization + Meta-optimization | L4 |
| Harnessing Agentic Evolution [237] | Improvement mechanism | Trace + Execution result + Outcome label + Improvement statistics | Meta-optimization | L4 |
| EvoX [240] | Improvement mechanism | Outcome label + Improvement statistics | Meta-optimization | L4 |
| ANCHOR [245] | Governance mechanism | Safety oversight | Governed commit | L3 |
| Statistical Gödel Machine [241] | Governance mechanism | Safety oversight | Governed commit | L3 |
| ANNEAL [243] | Governance mechanism | Safety oversight | Diagnostic repair + Governed commit | L3 |
3.1. Module-Level Self-Improvement
3.1.1. Memory Self-Modification
3.1.2. Skill and Tool Self-Improvement
3.1.3. Prompt and Context Self-Optimization
3.2. Orchestration and Architecture Search
3.2.1. Test-Time Compute and Trajectory Orchestration
3.2.2. Agentic Unit Composition
3.2.3. Execution Flow, Topology, and Workflow Optimization
3.3. Self-Referential Code Modification
3.3.1. Bounded Self-Harness and Policy Repair
3.3.2. Source-Code and Scaffold Self-Modification
3.3.3. Open-Ended and Governed Recursive Self-Modification
4. Agent Data System Self-Improvement
4.1. Self-Improvement in Agent Data Production
4.1.1. Environment Generation or Simulation
4.1.2. Task and Trajectory Synthesis
4.1.3. Verification and Quality Assurance
4.2. Self-Improvement in Agent Data Utilization
Rule-based Curriculum Adaptation.
Learned Curriculum Adaptation.
4.3. Multi-Module Co-Improvement in Agent Data System
4.3.1. Co-Improvement within Data Production
Co-Improvement of Synthesis and Verification.
Co-Improvement of Environment and Task Synthesis.
4.3.2. Co-Improvement of Data Production and Utilization
Capability Boundary Matching.
Goal-directed Curriculum Bridging.
| Work | Self-Improvement Process | Domain | Level | |
| Object | Evidence | |||
| Data Production | ||||
| EnvGen [246] | Environment generation | Success rate | Embodied agent | L3 |
| Adaptive Env Gen [247] | Environment generation | LLM analysis + Physical check | Embodied agent | L3 |
| ACCEL [248] | Environment generation | Positive value loss | Procedural RL | L3 |
| DRED [249] | Environment generation | Value loss scoring | Procedural RL | L3 |
| SimWorld Studio [250] | Environment generation | Rule + VLM | Embodied agent | L3 |
| WebRL [252] | Task synthesis | Critic score | Web agent | L3 |
| SeRL [254] | Task synthesis | Majority-voting + Difficulty | Math reasoning | L3 |
| Self-CriTeach [255] | Task spec + Trajectory | Planner check + Reward | Robotic planning | L3 |
| SAGE [256] | Task synthesis | Code execution / LLM judge | Reasoning | L3 |
| CoEvolve [253] | Task synthesis | Env. execution + Binary reward | Tool-use agent | L3 |
| SPIN [257] | Trajectory synthesis | Self-play | General | L3 |
| Arena Learning [258] | Trajectory synthesis | LLM judge | Dialogue | L3 |
| EVOLVE [259] | Trajectory synthesis | RM | General | L3 |
| DNPO [261] | Trajectory synthesis | LLM judge | General | L3 |
| PLD [262] | Trajectory synthesis | Binary reward + Success rate | Embodied agent | L3 |
| SI VLM Judges [263] | Judge model | Accuracy | Multimodal RM | L3 |
| Data Utilization | ||||
| AMC-TSI [267] | Curriculum | Accuracy + Transfer regret | Math reasoning | L3 |
| EvoCurr [265] | Curriculum + Memory | Rollout win rate | Game | L3 |
| TRUSTEE [264] | Curriculum | Eval pass rate | Tool-use agent | L3 |
| Actor-Curator [268] | Curriculum | Policy improvement | Math reasoning | L3 |
| Co-Improvement | ||||
| Agent0-VL [269] | Verifier + Trajectory | Self-verification | VL reasoning | L3 |
| ACE [270] | Adversary + Trajectory + Unit tests | Adversarial testing | Code generation | L3 |
| Agent-World [271] | Synthesized env/task | Rubrics/Validators | Tool-use agent | L3 |
| Agent0 [272] | Curriculum model + Tasks + Trajectory | Self-consistency + Majority vote | Reasoning | L3 |
| R-Zero [273] | Challenger model + Tasks + Trajectory | Majority-vote + Uncertainty | Reasoning | L3 |
5. Agent Trainer Self-Improvement
- Inner-Loop Trainer Adaptation operates within an active training lineage. Evidence from the current run directly changes persistent state in the supervision design, optimization strategy, or training infrastructure. The revised trainer state then governs subsequent model updates in the same lineage. The improvement mechanism remains fixed.
- Outer-Loop Trainer Search through Experiments operates across training experiments. Within each experiment, a trainer and a model interact through a training loop; across experiments, a fixed improvement mechanism uses the resulting evidence to propose runnable trainer candidates and evaluate them in bounded experiments. Based on the results, it rejects a candidate or integrates it as the retained trainer version for later experiments.
- Meta-Loop Improvement-Mechanism Evolution operates across successive trainer-improvement cycles. Within each cycle, the current improvement mechanism governs diagnosis, candidate proposal, evaluation, and integration. Across cycles, evidence from model-training experiments updates the persistent mechanism from to . The successor mechanism then governs these processes in later trainer-improvement cycles.
5.1. Inner-Loop Trainer Adaptation
5.1.1. Supervision Design Evolution
Evaluation Criteria Evolution
Evaluator and Verifier Evolution
Process-Level Credit Assignment Evolution
Training-Signal Routing and Internalization
5.1.2. Optimization Strategy Evolution
Objective Adaptation
Parameter-Update Adaptation
5.1.3. Training Infrastructure Evolution
5.2. Outer-Loop Trainer Search through Experiments
5.2.1. Experiment Diagnosis
Diagnostic Evidence Synthesis
Trainer Fault Localization
Diagnostic Hypothesis Validation
5.2.2. Trainer Candidate Search
Training-Configuration Search
Learning-Objective and Signal Search
Executable Training-Procedure Search
5.2.3. Trainer Candidate Evaluation
Full-Run Evaluation
Multi-Fidelity Evaluation
Robust Validation
5.3. Meta-Loop Improvement-Mechanism Evolution
5.3.1. Improvement-Model Evolution
5.3.2. Improvement-Harness Evolution
6. Co-Improvement of Multiple System Components
6.1. Harness-Trainer Co-Improvement
Training Conditioned on Agent Skills and Experience.
Harness Refinement Driven by Training Feedback.
6.2. Harness-Data Co-Improvement
Co-Improving Skills and Trajectory Generation.
Co-Improving Harness Code and Trajectory Generation.
6.3. Data-Trainer Co-Improvement
7. Open Problems and Future Research Directions
7.1. Long-Horizon Evaluation of Improvements in Real-World Production
7.2. Observable, Scalable, and Modifiable Training and Inference Infrastructure
7.3. From Bounded RSI to General RSI
7.4. Safety and Controllability Under Recursive Self-Modification
7.5. Human-Expert and Agent Co-Improvement
8. Conclusion
References
- METR. Measuring AI Ability to Complete Long Tasks. 2025. Available online: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/.
- Kwa, T.; West, B.; Becker, J.; Deng, A.; Garcia, K.; Hasin, M.; Jawhar, S.; Kinniment, M.; Rush, N.; Arx, S.V.; et al. Measuring AI Ability to Complete Long Software Tasks, 2026. arXiv arXiv:cs.
- Qwen Team. Qwen3.7: The Agent Frontier. 2026. [Google Scholar] [CrossRef]
- Xu, A.; Lin, B.; Xue, B.; Wang, B.; Xu, B.; Wu, B.; Zhang, B.; Lin, C.; Dong, C.; Ling, C.; et al. Deepseek-v4: Towards highly efficient million-token context intelligence. arXiv 2026, arXiv:2606.19348. [Google Scholar]
- Blakeman, A.; Thomas, A.; Jhunjhunwala, A.; Gupta, A.; Khattar, A.; Rajfer, A.; Renduchintala, A.; Asif, A.; Vavre, A.; Miranda, A.F.; et al. Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning. arXiv 2026, arXiv:2606.15007. [Google Scholar]
- Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv 2024, arXiv:cs. [Google Scholar]
- Deng, X.; Da, J.; Pan, E.; He, Y.Y.; Ide, C.; Garg, K.; Lauffer, N.; Park, A.; Pasari, N.; Rane, C.; et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv 2025, arXiv:cs. [Google Scholar]
- Merrill, M.A.; Shaw, A.G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J.Y.; Walshe, T.; Buchanan, E.K.; et al. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. arXiv 2026, arXiv:cs. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv 2023, arXiv:cs. [Google Scholar]
- Yang, J.; Jimenez, C.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O. Swe-agent: Agent-computer interfaces enable automated software engineering. Adv. Neural Inf. Process. Syst. 2024, 37, 50528–50652. [Google Scholar] [CrossRef]
- Anthropic. Claude Code AI coding agent harness. 2025. Available online: https://docs.anthropic.com/en/docs/claude-code.
- OpenAI. OpenAI Codex CLI Lightweight coding agent running in the terminal. 2025. Available online: https://github.com/openai/codex.
- Blakeman, A.; Grattafiori, A.; Basant, A.; Gupta, A.; Khattar, A.; Renduchintala, A.; Vavre, A.; Shukla, A.; Bercovich, A.; Ficek, A.; et al. NVIDIA Nemotron 3: Efficient and Open Intelligence. arXiv 2025, arXiv:2512.20856. [Google Scholar]
- Zhao, J.; Chen, G.; Meng, F.; Li, M.; Chen, J.; Xu, H.; Sun, Y.; Zhao, W.X.; Song, R.; Zhang, Y.; et al. Immersion in the github universe: Scaling coding agents to mastery. arXiv 2026, arXiv:2602.09892. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Kunc, Y.; et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv 2024, arXiv:2402.03300. [Google Scholar]
- Sheng, G.; Zhang, C.; Ye, Z.; Wu, X.; Zhang, W.; Zhang, R.; Peng, Y.; Lin, H.; Wu, C. HybridFlow: A Flexible and Efficient RLHF Framework. arXiv 2409.19256. 2024. [Google Scholar]
- Zhu, Z.; Xie, C.; Lv, X. slime Contributors. slime: An LLM post-training framework for RL Scaling GitHub repository. Corresponding author; Lv, Xin, Ed.; 2025; Available online: https://github.com/THUDM/slime.
- Zhang, L.; Chen, M.; Cao, R.; Chen, J.; Zhou, F.; Xu, Y.; Yang, J.; Ma, Z.; Chen, L.; Luo, C.; et al. MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era, 2026. arXiv arXiv:cs.
- Wu, R.; Wang, X.; Mei, J.; Cai, P.; Fu, D.; Yang, C.; Wen, L.; Yang, X.; Shen, Y.; Wang, Y.; et al. EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle. arXiv 2025, arXiv:cs. [Google Scholar]
- Novikov, A.; Vu, N.; Eisenberger, M.; Dupont, E.; Huang, P.S.; Wagner, A.Z.; Shirobokov, S.; Kozlovskii, B.; Ruiz, F.J.R.; Mehrabian, A.; et al. AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery. arXiv 2025, arXiv:cs. [Google Scholar]
- Lin, J.; Liu, S.; Pan, C.; Lin, L.; Dou, S.; Huang, X.; Yan, H.; Han, Z.; Gui, T. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses. CoRR 2026, abs/2604.25850, [2604.25850. [Google Scholar] [CrossRef]
- Zhang, J.; Zhao, B.; Yang, W.; Foerster, J.N.; Clune, J.; Jiang, M.; Devlin, S.; Shavrina, T. Hyperagents. abs/2603.19461; CoRR. 2026; p. 2603.19461. [Google Scholar] [CrossRef]
- Good, I.J. Speculations Concerning the First Ultraintelligent Machine. In Advances in Computers; Elsevier, 1966; Volume 6, pp. 31–88. [Google Scholar] [CrossRef]
- Schmidhuber, J. Gödel Machines: Fully Self-referential Optimal Universal Self-improvers. In Artificial General Intelligence; Cognitive Technologies; Goertzel, B., Pennachin, C., Eds.; Springer, 2007; pp. 199–226. [Google Scholar] [CrossRef]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: language agents with verbal reinforcement learning. In Proceedings of the Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023; New Orleans, LA, USA, Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 10 - 16 December 2023. [Google Scholar]
- Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.; Huang, G. ExpeL: LLM Agents Are Experiential Learners. CoRR 2023, abs/2308.10144, [2308.10144. [Google Scholar] [CrossRef]
- Chen, S.; Lin, S.; Gu, X.; Shi, Y.; Lian, H.; Yun, L.; Chen, D.; Sun, W.; Cao, L.; Wang, Q. SWE-Exp: Experience-Driven Software Issue Resolution. CoRR 2025, abs/2507.23361, [2507.23361. [Google Scholar] [CrossRef]
- Ouyang, S.; Yan, J.; Hsu, I.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L.T.; Daruki, S.; Tang, X.; et al. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. CoRR 2025, abs/2509.25140, [2509.25140. [Google Scholar] [CrossRef]
- Cai, Z.; Guo, X.; Pei, Y.; Feng, J.; Chen, J.; Zhang, Y.; Ma, W.; Wang, M.; Zhou, H. FLEX: Continuous Agent Evolution via Forward Learning from Experience. CoRR 2025, abs/2511.06449, [2511.06449. [Google Scholar] [CrossRef]
- Shen, M.; Zha, K.; He, Z.; Hong, Z.; Ouyang, S.; Ryu, J.J.; Sattigeri, P.; Diggavi, S.N.; Wornell, G.W. Decocted Experience Improves Test-Time Inference in LLM Agents. abs/2604.04373; CoRR. 2026; p. 2604.04373. [Google Scholar] [CrossRef]
- Packer, C.; Fang, V.; Patil, S.G.; Lin, K.; Wooders, S.; Gonzalez, J.E. MemGPT: Towards LLMs as Operating Systems. CoRR 2023, abs/2310.08560, [2310.08560. [Google Scholar] [CrossRef]
- Xu, W.; Liang, Z.; Mei, K.; Gao, H.; Tan, J.; Zhang, Y. A-MEM: Agentic Memory for LLM Agents. CoRR 2025, abs/2502.12110, [2502.12110. [Google Scholar] [CrossRef]
- Liu, J.; Su, Y.; Xia, P.; Han, S.; Zheng, Z.; Xie, C.; Ding, M.; Yao, H. SimpleMem: Efficient Lifelong Memory for LLM Agents. CoRR 2026, abs/2601.02553, [2601.02553. [Google Scholar] [CrossRef]
- Dai, Z.; Deng, S.; Guan, S.; Tian, Y.; Yao, X.; Yan, X.; Cheng, J. RecMem: Recurrence-based Memory Consolidation for Efficient and Effective Long-Running LLM Agents. CoRR 2026, abs/2605.16045, [2605.16045. [Google Scholar] [CrossRef]
- Jin, Y.; Zhang, S.; Wang, H.; Qin, L.; Zhang, Y.; Zhang, W. EXG: Self-Evolving Agents with Experience Graphs. CoRR 2026, abs/2605.17721, [2605.17721. [Google Scholar] [CrossRef]
- Fang, J.; Xu, B.; Wang, Z.; Cao, H.; Deng, X.; Dong, B.; Zhu, H.; Huang, R.; Yu, G.; Wei, Y.; et al. Rethinking Memory as Continuously Evolving Connectivity. abs/2605.28773; CoRR. 2026; p. 2605.28773. [Google Scholar] [CrossRef]
- Ji, S.; Wu, B.; Wang, Z.; Xia, L.; Li, Q.; Wang, R.; Ding, W.; Zhu, Z.; Li, B.; Dai, G.; et al. Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory. 2026. [Google Scholar]
- Fei, T.; Song, M.; Zheng, M.; Yu, X. Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents. 2026. [Google Scholar]
- Ye, C.; Liu, Y.; Wang, Y.; Yu, H.; Zhao, Y.; Liu, G.; McAuley, J.J.; You, J. Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents. CoRR 2026, abs/2605.20616, [2605.20616. [Google Scholar] [CrossRef]
- Zhang, Y.; Wu, Y.; Yu, Y.; Wu, Q.; Wang, H. Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback. CoRR 2026, abs/2602.02369, [2602.02369. [Google Scholar] [CrossRef]
- Liao, J.; Shi, H.; Zhou, R.; Wang, J.; Zhang, S.; Zhang, W.; Wen, Y.; Li, Z.; Xiong, F.; Tang, B.; et al. MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs. CoRR 2026, abs/2605.08374, [2605.08374. [Google Scholar] [CrossRef]
- Wang, X.; Mao, W.; Wu, J.; Wang, X.; He, X. R2-Mem: Reflective Experience for Memory Search. CoRR 2026, abs/2605.13486, [2605.13486. [Google Scholar] [CrossRef]
- Zhang, Y.; Li, Y.; Payani, A.; Wang, L. AdaMEM: Test-Time Adaptive Memory for Language Agents. 2026. [Google Scholar]
- Kim, K.; Kang, M.; Kim, T.; Yang, Y.; Ren, M.; Hwang, S.J. Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents. CoRR 2026, abs/2604.14004, [2604.14004. [Google Scholar] [CrossRef]
- Cheng, Y.; Zhou, J.; Hu, Y.; Chen, Y.; Zhou, H.; Chen, M.; Zhang, Z.; Shao, K.; Xie, Y.; Yin, Z. TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking. CoRR 2026, abs/2602.03224, [2602.03224. [Google Scholar] [CrossRef]
- Wang, S.; Brahma, D.; Henao, R. SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs. CoRR 2026, abs/2605.30711, [2605.30711. [Google Scholar] [CrossRef]
- Song, Y.; Xin, Q. D-MEM: Dopamine-Gated Agentic Memory via Reward Prediction Error Routing. CoRR 2026, abs/2603.14597, [2603.14597. [Google Scholar] [CrossRef]
- Zhang, G.; Ren, H.; Zhan, C.; Zhou, Z.; Wang, J.; Zhu, H.; Zhou, W.; Yan, S. MemEvolve: Meta-Evolution of Agent Memory Systems. CoRR 2025, abs/2512.18746, [2512.18746. [Google Scholar] [CrossRef]
- Xiong, Y.; Hu, S.; Clune, J. Learning to Continually Learn via Meta-learning Agentic Memory Designs. abs/2602.07755; CoRR. 2026; p. 2602.07755. [Google Scholar] [CrossRef]
- Pan, W.; Liu, S.; Zhou, X.; Zhang, S.; Shi, W.; Xu, M.; Jia, X. M*: Every Task Deserves Its Own Memory Harness. CoRR 2026, abs/2604.11811, [2604.11811. [Google Scholar] [CrossRef]
- Liu, J.; Ye, X.; Xia, P.; Zheng, Z.; Xie, C.; Ding, M.; Yao, H. EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents. CoRR 2026, abs/2605.13941, [2605.13941. [Google Scholar] [CrossRef]
- Liu, Q.; Wang, G.; Wu, W.; Huang, J.; Tao, X.; Song, D.; Zhou, J.; He, L. MemPro: Agentic Memory Systems as Evolvable Programs. 2026. [Google Scholar]
- Zhang, H.; Long, Q.; Bao, J.; Feng, T.; Zhang, W.; Yue, H.; Wang, W. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents. CoRR 2026, abs/2602.02474, [2602.02474. [Google Scholar] [CrossRef]
- Yang, Y.; Liu, T.; Zhu, W.B.; Shi, T.; Song, L.; Jia, R. Self-Evolving LLM Memory Extraction Across Heterogeneous Tasks. CoRR 2026, abs/2604.11610, 2604.11610. [Google Scholar] [CrossRef]
- Qian, C.; Han, C.; Fung, Y.R.; Qin, Y.; Liu, Z.; Ji, H. CREATOR: Disentangling Abstract and Concrete Reasonings of Large Language Models through Tool Creation. CoRR 2023, abs/2305.14318, [2305.14318. [Google Scholar] [CrossRef]
- Cai, T.; Wang, X.; Ma, T.; Chen, X.; Zhou, D. Large Language Models as Tool Makers. CoRR 2023, abs/2305.17126, [2305.17126. [Google Scholar] [CrossRef]
- Ding, H.; Tao, S.; Pang, L.; Wei, Z.; Gao, J.; Ding, B.; Shen, H.; Cheng, X. ToolCoder: A Systematic Code-Empowered Tool Learning Framework for Large Language Models. CoRR 2025, abs/2502.11404, [2502.11404. [Google Scholar] [CrossRef]
- Liu, X.; Yin, D.; Wu, Z.; Feng, Y. RefTool: Enhancing Model Reasoning with Reference-Guided Tool Creation. CoRR 2025, abs/2505.21413, [2505.21413. [Google Scholar] [CrossRef]
- Zheng, B.; Fatemi, M.Y.; Jin, X.; Wang, Z.Z.; Gandhi, A.; Song, Y.; Gu, Y.; Srinivasa, J.; Liu, G.; Neubig, G.; et al. SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills. CoRR 2025, abs/2504.07079, [2504.07079. [Google Scholar] [CrossRef]
- Wang, C.; Yu, Z.; Xie, X.; Yao, W.; Fang, R.; Qiao, S.; Cao, K.; Zheng, G.; Qi, X.; Zhang, P.; et al. SkillX: Automatically Constructing Skill Knowledge Bases for Agents. CoRR 2026, abs/2604.04804, 2604.04804. [Google Scholar] [CrossRef]
- Zhang, Y.; Han, X.; Jiang, X.; Wang, R. Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition. 2026. [Google Scholar] [CrossRef] [PubMed]
- Xiao, C.; Jiao, Z.; Wang, S.; Wang, W.; Zhao, B.; Wei, H.; Zhang, L.; Qu, L. Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills. 2026. [Google Scholar] [CrossRef]
- Pan, Q.; Yang, Y.; Li, J.; Zhou, J.; Chen, K.; Li, X.; Chen, Q.; He, L. Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents. 2026. [Google Scholar] [CrossRef]
- Yan, Z.; Song, D.; Zhang, H.; Liang, W.; Zhang, Y.; Dai, Y.; He, L.; Yu, P.S.; Xu, R.; Li, X.; et al. OpenSkill: Open-World Self-Evolution for LLM Agents. 2026. [Google Scholar] [CrossRef]
- Shen, S.; Cheng, W.; Ma, M.; Turcan, A.; Zhang, M.J.; Ma, J. SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources. CoRR 2026, abs/2604.03964, [2604.03964. [Google Scholar] [CrossRef]
- Qiu, Z.; Song, K.; Tang, S.; Qiao, S.; Liang, L.; Chen, H.; Deng, S. Unsupervised Skill Discovery for Agentic Data Analysis. 2026. [Google Scholar] [CrossRef]
- Lin, H.; Li, P.; Song, J.; Jiang, F.; Zhang, T. MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation. CoRR 2026, abs/2605.27366, [2605.27366. [Google Scholar] [CrossRef]
- Ouyang, S.; Yan, J.; Chen, Y.; Han, R.; Wang, Z.; Mishra, B.D.; Meng, R.; Li, C.; Jiao, Y.; Zha, K.; et al. SkillOS: Learning Skill Curation for Self-Evolving Agents. CoRR 2026, abs/2605.06614, [2605.06614. [Google Scholar] [CrossRef]
- Ma, Z.; Yang, S.; Ji, Y.; Wang, X.; Wang, Y.; Hu, Y.; Huang, T.; Chu, X. SkillClaw: Let Skills Evolve Collectively with Agentic Evolver. CoRR 2026, abs/2604.08377, [2604.08377. [Google Scholar] [CrossRef]
- Yue, M.; Liu, Z.; Yang, L.; Zhang, J.; Liu, Z.; Chen, H.; Yao, Z.; Savarese, S.; Xiong, C.; Heinecke, S.; et al. ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning. CoRR 2025, abs/2510.07768, [2510.07768. [Google Scholar] [CrossRef]
- Yang, Y.; Gong, Z.; Huang, W.; Yang, Q.; Zhou, Z.; Huang, Z.; Li, Y.; Gao, X.; Dai, Q.; Liu, B.; et al. SkillOpt: Executive Strategy for Self-Evolving Agent Skills. CoRR 2026, abs/2605.23904, [2605.23904. [Google Scholar] [CrossRef]
- Wang, Y.; Zhou, Y.; Liang, Y.; Zhang, C.; Liu, F.; Zhou, J.; Yao, H. Not All Skills Help: Measuring and Repairing Agent Knowledge. 2026. [Google Scholar] [CrossRef]
- Zhang, Q.; Feng, Z.; Shi, X.; Hu, X.; Liu, C.; Xie, P.; Wang, X.; Ye, J.; Hooi, B.; Wang, H.; et al. SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization. 2026. [Google Scholar]
- Shi, H.; Yuan, X.; Liu, B. Evolving Programmatic Skill Networks. abs/2601.03509; CoRR. 2026; p. 2601.03509. [Google Scholar] [CrossRef]
- Gautam, S.; Radhakrishna, A.; Gulwani, S. SkillAxe: Sharpening LLM-Authored Agent Skills Through Evaluation-Guided Self-Refinement. 2026. [Google Scholar]
- Gao, H.; Chen, H.; Wang, C.; Guo, S.; Pang, L.; Liu, Z.; Shen, H.; Cheng, X. SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing. 2026. [Google Scholar] [CrossRef]
- Ma, Y.; Huang, Y.; Bao, H.; Zhuang, H.; Shukla, S.; Galley, M.; Zhang, X.; Feuerriegel, S. SkillGen: Verified Inference-Time Agent Skill Synthesis. CoRR 2026, abs/2605.10999, [2605.10999. [Google Scholar] [CrossRef]
- Zhang, H.; Fan, S.; Zou, H.P.; Chen, Y.; Wang, Z.; Zhou, J.; Li, C.; Huang, W.; Yao, Y.; Zheng, K.; et al. CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification. abs/2604.01687; CoRR. 2026; p. 2604.01687. [Google Scholar] [CrossRef]
- Yang, Y.; Bhatt, N.P.; Wang, K.; Tetteh, S.; Wang, Z.; Topcu, U. VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents. 2026. [Google Scholar]
- Lu, J.; Kong, Z.; Wang, Y.; Fu, R.; Wan, H.; Yang, C.; Lou, W.; Sun, H.; Wang, L.; Jiang, Y.; et al. Beyond Static Tools: Test-Time Tool Evolution for Scientific Reasoning. CoRR 2026, abs/2601.07641, [2601.07641. [Google Scholar] [CrossRef]
- Wei, S.; Min, H.S.; Dong, X.; Lin, X.; Cui, S.; Jiang, B.; Dai, Z.; Kuang, K.; Xu, G.; Wu, F.; et al. MetaForge: A Self-Evolving Multimodal Agent that Retrieves, Adapts, and Forges Tools On Demand. 2026. [Google Scholar]
- Wei, Y.; Huang, Z.; Lu, S.; Qian, J.; Wang, Q.; Wu, C.; He, L. SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems. 2026. [Google Scholar]
- Wang, J.; Yan, Q.; Wang, Y.; Tian, Y.; Mishra, S.S.; Xu, Z.; Gandhi, M.; Xu, P.; Cheong, L.L. Reinforcement Learning for Self-Improving Agent with Skill Library. CoRR 2025, abs/2512.17102, [2512.17102. [Google Scholar] [CrossRef]
- Wong, S.; Qi, Z.; Wang, Z.; Hu, N.; Lin, S.; Ge, J.; Gao, E.; Chen, W.; Du, Y.; Yu, M.; et al. Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases. CoRR 2025, abs/2512.10398, [2512.10398. [Google Scholar] [CrossRef]
- Yang, C.; Wang, X.; Lu, Y.; Liu, H.; Le, Q.V.; Zhou, D.; Chen, X. Large Language Models as Optimizers. CoRR 2023, abs/2309.03409, 2309.03409. [Google Scholar] [CrossRef]
- Fernando, C.; Banarse, D.; Michalewski, H.; Osindero, S.; Rocktäschel, T. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. CoRR 2023, abs/2309.16797, [2309.16797. [Google Scholar] [CrossRef]
- Xiang, J.; Zhang, J.; Yu, Z.; Teng, F.; Tu, J.; Liang, X.; Hong, S.; Wu, C.; Luo, Y. Self-Supervised Prompt Optimization. CoRR 2025, abs/2502.06855, [2502.06855. [Google Scholar] [CrossRef]
- Peng, D.; Zhou, Y.; Chen, Q.; Liu, J.; Chen, J.; Qin, L. DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective. CoRR 2025, abs/2503.13413, 2503.13413. [Google Scholar] [CrossRef]
- Lin, Z.J.; Letham, B.; Dooley, S.; Balandat, M.; Bakshy, E. Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts. CoRR 2026, abs/2605.19093, [2605.19093. [Google Scholar] [CrossRef]
- Singhal, R.; Tambwekar, P.; Maamari, K. PrefPO: Pairwise Preference Prompt Optimization. abs/2603.19311; CoRR. 2026; p. 2603.19311. [Google Scholar] [CrossRef]
- Wang, F.; Si, S.; Hsieh, C.J.; Dhillon, I.S. APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection. 2026. [Google Scholar]
- Kang, E.H.; Yoganarasimhan, H. Bayesian Optimization in Language Space: An Eval-Efficient AI Self-Improvement Framework. CoRR 2025, abs/2511.12063, [2511.12063. [Google Scholar] [CrossRef]
- Lee, Y.; Boen, J.; Finn, C. Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison. ArXiv 2025, abs/2511.07919. [Google Scholar]
- Lu, M.; Feng, C.; Han, H.; Lu, G.; Sun, Y.; Ding, X.; Long, S.; Li, F.; Motwani, T. SPEAR: Code-Augmented Agentic Prompt Optimization. CoRR 2026, abs/2605.26275, [2605.26275. [Google Scholar] [CrossRef]
- Fernandes, R.C.; Fehring, L.; Eimer, T.; Lindauer, M.; Feurer, M. Environment-Grounded Automated Prompt Optimization for LLM Game Agents. arXiv 2026, arXiv:2606.17838. [Google Scholar]
- Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Vardhamanan, S.; Haq, S.; Sharma, A.; Joshi, T.T.; Moazam, H.; et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. CoRR 2023, abs/2310.03714, [2310.03714. [Google Scholar] [CrossRef]
- Opsahl-Ong, K.; Ryan, M.J.; Purtell, J.; Broman, D.; Potts, C.; Zaharia, M.; Khattab, O. Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs. CoRR 2024, abs/2406.11695, [2406.11695. [Google Scholar] [CrossRef]
- Spiess, C.; Vaziri, M.; Mandel, L.; Hirzel, M. AutoPDL: Automatic Prompt Optimization for LLM Agents. CoRR 2025, abs/2504.04365, 2504.04365. [Google Scholar] [CrossRef]
- Lin, S.; Hua, W.; Li, L.; Wang, Z.; Zhang, Y. ADO: Automatic Data Optimization for Inputs in LLM Prompts. CoRR 2025, abs/2502.11436, [2502.11436. [Google Scholar] [CrossRef]
- Shankaranarayanan, A.; Venkataraman, A.N.; Nikolakopoulos; Kumaraswamy, V.S.; Zhang, T.; Chander, S.; Saboo, R.R.; Khan, S.A. A FRAMEWORK FOR PROMPT OPTIMIZATION AND TRANSLATION A CROSS FOUNDATION MODELS. [CrossRef]
- Yüksekgönül, M.; Bianchi, F.; Boen, J.; Liu, S.; Huang, Z.; Guestrin, C.; Zou, J. TextGrad: Automatic "Differentiation" via Text. abs/2406.07496; CoRR. 2024; p. 2406.07496. [Google Scholar] [CrossRef]
- Agrawal, L.A.; Tan, S.; Soylu, D.; Ziems, N.; Khare, R.; Opsahl-Ong, K.; Singhvi, A.; Shandilya, H.; Ryan, M.J.; Jiang, M.; et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. ArXiv 2025, abs/2507.19457. [Google Scholar]
- Chen, M.; Deng, W.; Zou, J.; Yu, H.; Li, X. Textual Equilibrium Propagation for Deep Compound AI Systems. CoRR 2026, abs/2601.21064, [2601.21064. [Google Scholar] [CrossRef]
- Rishav, R.; Pujari, P.; Rastogi, P. ContraPrompt: Contrastive Prompt Optimization via Dyadic Reasoning Trace Analysis. CoRR 2026, abs/2604.17937, [2604.17937. [Google Scholar] [CrossRef]
- Zhu, T.; Yao, T.; Kuwaranancharoen, K.; Singh, A.; Lai, Y.; Mohan, D.A.; Bhargava, S. Graph-based Target Back-Propagation for Context Adaptation in Multi-LLM Agentic Systems. 2026. [Google Scholar]
- Li, W.; Song, Y.; Zhao, M.; Jin, B.; Li, W. Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization. CoRR 2026, abs/2605.30227, [2605.30227. [Google Scholar] [CrossRef]
- Xu, W.; Liu, S.; Wang, M. EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents. 2026. [Google Scholar]
- Li, H.; He, R.; Zhang, Q.; Ji, C.; Mang, Q.; Chen, X.; Agrawal, L.A.; Liao, W.; Yang, E.; Cheung, A.; et al. Combee: Scaling Prompt Learning for Self-Improving Language Model Agents. CoRR 2026, abs/2604.04247, [2604.04247. [Google Scholar] [CrossRef]
- Suzgun, M.; Yüksekgönül, M.; Bianchi, F.; Jurafsky, D.; Zou, J. Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory. CoRR 2025, abs/2504.07952, [2504.07952. [Google Scholar] [CrossRef]
- Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; et al. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. CoRR 2025, abs/2510.04618, [2510.04618. [Google Scholar] [CrossRef]
- Pei, Z.; Zhen, H.; Kai, S.; Pan, S.J.; Wang, Y.; Yuan, M.; Yu, B. SCOPE: Prompt Evolution for Enhancing Agent Effectiveness. CoRR 2025, abs/2512.15374, [2512.15374. [Google Scholar] [CrossRef]
- Vassilyev, N.; Berrios, W.; Zhang, R.; Han, B.; Kiela, D.; Mehri, S. Reflective Context Learning: Studying the Optimization Primitives of Context Space. abs/2604.03189; CoRR. 2026; p. 2604.03189. [Google Scholar] [CrossRef]
- Zhu, Z.; Hu, Y.; Dai, Y.; Fang, J.; Jiang, C.; Hu, S.; Zhao, Y. Unified Context Evolution for LLM Agents. 2026. [Google Scholar] [CrossRef]
- Parashar, J.; Bhandarkar, S. KACE: Knowledge-Adaptive Context Engineering for Mathematical Reasoning; 2026. [Google Scholar]
- Huang, Z.; Kuncoro, A.; Feng, Q.; Shen, J.; Dery, L.M.; Szlam, A.; Ranzato, M. Context Training with Active Information Seeking. CoRR 2026, abs/2605.13050, [2605.13050. [Google Scholar] [CrossRef]
- Xie, Y.; Wang, K.; Cheng, B.; Yao, J.; Sha, Z.; Duffy, A.; Xi, Y.; Mei, H.; Tan, C.; Wei, C.; et al. MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games. CoRR 2026, abs/2603.09022, [2603.09022. [Google Scholar] [CrossRef]
- Wu, Y.; Long, W.; Nguyen, C.T.; Wang, X.; et al. Contrastive Self-Refinement for Low-Cost Adaptation in Real-World Text-to-SQL. [CrossRef]
- Zha, J.; Wang, J.; Zhou, C.; Song, X. Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents. 2026. [Google Scholar] [CrossRef]
- Tao, W.; Wu, H.; Wong, W.F. SePO: Self-Evolving Prompt Agent for System Prompt Optimization. 2026. [Google Scholar] [CrossRef]
- Chen, X.; Xu, C.; Wang, Y.; Liu, B.; Yao, Z.; He, Y. Learning to Self-Evolve. CoRR 2026, abs/2603.18620, [2603.18620. [Google Scholar] [CrossRef]
- Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.L.; Cao, Y.; Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. CoRR 2023, abs/2305.10601, [2305.10601. [Google Scholar] [CrossRef]
- Zhou, A.; Yan, K.; Shlapentokh-Rothman, M.; Wang, H.; Wang, Y. Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models. CoRR 2023, abs/2310.04406, [2310.04406. [Google Scholar] [CrossRef]
- Zhu, K.; Li, H.; Wu, S.; Xing, T.; Ma, D.; Tang, X.; Liu, M.; Yang, J.; Liu, J.; Jiang, Y.E.; et al. Scaling Test-time Compute for LLM Agents. CoRR 2025, abs/2506.12928, [2506.12928. [Google Scholar] [CrossRef]
- Antoniades, A.; Örwall, A.; Zhang, K.; Xie, Y.; Goyal, A.; Wang, W.Y. SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement. CoRR 2024, abs/2410.20285, [2410.20285. [Google Scholar] [CrossRef]
- Kim, J.; Yang, W.; Niu, K.; Zhang, H.; Zhu, Y.; Helenowski, E.; Silva, R.; Chen, Z.; Iyer, S.; Zaheer, M.; et al. Scaling Test-Time Compute for Agentic Coding. CoRR 2026, abs/2604.16529, [2604.16529. [Google Scholar] [CrossRef]
- Zeng, G.; Shen, M.; Chen, D.; Qi, Z.; Das, S.; Gutfreund, D.; Cox, D.; Wornell, G.W.; Lu, W.; Hong, Z.; et al. Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering. CoRR 2025, abs/2505.23604, [2505.23604. [Google Scholar] [CrossRef]
- Lin, J.; Guo, Y.; Han, Y.; Hu, S.; Ni, Z.; Wang, L.; Chen, M.; Liu, H.; Chen, R.; He, Y.; et al. SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents. CoRR 2025, abs/2508.02085, [2508.02085. [Google Scholar] [CrossRef]
- Tan, D.Y.Y.; Chin, K.; Zhang, J. AgentGA: Evolving Code Solutions in Agent-Seed Space. CoRR 2026, abs/2604.14655, [2604.14655. [Google Scholar] [CrossRef]
- Fattha, A.D.; Chua, K.Y.; Jiang, L.; Wynter, L. Exploration Structure in LLM Agents for Multi-File Change Localization. 2026. [Google Scholar]
- Zhang, S.; Wang, M.; Shi, Y.; Wang, Y.; Gu, X.; Yao, Y.; Fu, R.; Fu, S. FastContext: Training Efficient Repository Explorer for Coding Agents. 2026. [Google Scholar]
- Han, H.; Xie, J.; Ma, X.; Zhu, W.; Zhang, Z.; Long, Z.; Chen, H.; Ye, Q. SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling. CoRR 2026, abs/2604.14820, [2604.14820. [Google Scholar] [CrossRef]
- Mao, C.; Lei, Y.; Wei, Z.; Liang, M.; Wang, Z.; Xu, J.; Chen, D.; Jiang, W.; Li, Y. EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering. CoRR 2026, abs/2602.05242, [2602.05242. [Google Scholar] [CrossRef]
- Tan, B.; Deng, H.; Zhang, J.; Xu, J.; He, P.; Sun, Y. SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding. CoRR 2026, abs/2601.22956, [2601.22956. [Google Scholar] [CrossRef]
- Liu, M.; Chen, Z.; Pei, Z.; Wang, Z.; Wang, Y.; Zheng, Z. Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition. CoRR 2026, abs/2603.01814, [2603.01814. [Google Scholar] [CrossRef]
- Ding, Y.; Zhang, L. SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents. abs/2601.22129; CoRR. 2026; p. 2601.22129. [Google Scholar] [CrossRef]
- Zhang, T.; Popa, A.; Xu, Y.; Song, R.; Dimitriadis, D. PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement. CoRR 2026, abs/2605.11225, [2605.11225. [Google Scholar] [CrossRef]
- Pan, R.; Wang, J.; Zhang, Q.; Zhu, Y.; Wu, L.; Yang, Z.; Zhang, Y.; Zhang, L.; Zhang, H. Persistent Cross-Attempt State Optimization for Repository-Level Code Generation. CoRR 2026, abs/2604.03632, [2604.03632. [Google Scholar] [CrossRef]
- Zhou, Z.; Cao, C.; Feng, X.; Li, X.; Li, Z.; Lu, X.; Yao, J.; Huang, W.; Cheng, T.; Zhang, J.; et al. AlphaApollo: A System for Deep Agentic Reasoning. 2025. [Google Scholar] [CrossRef]
- Chen, P.B.; Zhang, Y.; Roth, D.; Madden, S.; Andreas, J.; Cafarella, M.J. Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation. CoRR 2025, abs/2505.14398, [2505.14398. [Google Scholar] [CrossRef]
- Tran, Q.M.; Huang, Z.; Zhang, W.; Han, B.; Yatani, K.; Sugiyama, M.; Liu, T. Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents. CoRR 2026, abs/2602.05810, [2602.05810. [Google Scholar] [CrossRef]
- Yang, X.; Zhou, J.; Pacheco, M.; Zhu, W.; He, P.; Wang, S.; Liu, K.; Pan, R. Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling. CoRR 2025, abs/2510.11838, [2510.11838. [Google Scholar] [CrossRef]
- Ma, R.; Jiang, Y.; Zhang, S.; Ma, Z.; Feng, Y.; Ng, V.; Wang, Z.; Yue, X.; Li, C.; Lu, L. FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair. CoRR 2026, abs/2603.17826, [2603.17826. [Google Scholar] [CrossRef]
- Hu, H.; Xie, G.; Zhang, Q.; Liu, J.; Yu, S.; Fang, C.; Chen, Z.; Xiao, L. EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution. CoRR 2026, abs/2605.30105, [2605.30105. [Google Scholar] [CrossRef]
- Yu, S.; Chong, D.; Nandi, A.; Soylu, D.; Sun, J.; Manning, C.D.; Shi, W. Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace. CoRR 2026, abs/2605.10913, [2605.10913. [Google Scholar] [CrossRef]
- Dong, Y.; He, J.; Hou, Y.; Du, D.; Xu, Z.; Yu, S.; Xia, Y.; Chen, H. DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback. CoRR 2026, abs/2605.22781, [2605.22781. [Google Scholar] [CrossRef]
- Zhang, X.; Wang, D.; Xu, K.; Zhu, Q.; Che, W. Scaling Laws for Agent Harnesses via Effective Feedback Compute. CoRR 2026, abs/2605.29682, [2605.29682. [Google Scholar] [CrossRef]
- Shang, Y.; Li, Y.; Zhao, K.; Ma, L.; Liu, J.; Xu, F.; Li, Y. AgentSquare: Automatic LLM Agent Search in Modular Design Space. CoRR 2024, abs/2410.06153, [2410.06153. [Google Scholar] [CrossRef]
- Li, Y.; Li, L.; Wu, Z.; Liao, Q.; Hao, J.; Shao, K.; Xu, F.; Li, Y. AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search. CoRR 2025, abs/2506.06017, [2506.06017. [Google Scholar] [CrossRef]
- Chen, J.; Shen, J.; Kang, H.; Hong, Z.; Jiang, Q.; Bose, S.; Zhang, Y.; Leng, L.; Vyas, A.K.; Mao, L.; et al. AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition. 2026. [Google Scholar] [CrossRef]
- Chen, G.; Dong, S.; Shu, Y.; Zhang, G.; Sesay, J.; Karlsson, B.F.; Fu, J.; Shi, Y. AutoAgents: A Framework for Automatic Agent Generation. CoRR 2023, abs/2309.17288, [2309.17288. [Google Scholar] [CrossRef]
- Yuan, S.; Song, K.; Chen, J.; Tan, X.; Li, D.; Yang, D. EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms. CoRR 2024, abs/2406.14228, [2406.14228. [Google Scholar] [CrossRef]
- Hu, S.; Lu, C.; Clune, J. Automated Design of Agentic Systems. CoRR 2024, abs/2408.08435, [2408.08435. [Google Scholar] [CrossRef]
- Ye, R.; Tang, S.; Ge, R.; Du, Y.; Yin, Z.; Chen, S.; Shao, J. MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems. CoRR 2025, abs/2503.03686, [2503.03686. [Google Scholar] [CrossRef]
- Xu, A.; Tai, Y. Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems. CoRR 2026, abs/2605.25233, [2605.25233. [Google Scholar] [CrossRef]
- Chen, X.; Liu, Y.; Wei, H.; Ding, K. LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning. CoRR 2026, abs/2605.14483, [2605.14483. [Google Scholar] [CrossRef]
- Feng, Y.; Luo, H.; Lin, Z.; Sun, Y.; Wei, P.; Hsieh, L.B.; Luu, A.T. OrchMAS: Orchestrated Reasoning with Multi Collaborative Heterogeneous Scientific Expert Structured Agents. CoRR 2026, abs/2603.03005, [2603.03005. [Google Scholar] [CrossRef]
- Hu, Y.; Zhang, Y.; Trager, M.; Zhang, Y.E.; Yang, S.; Xia, W.; Soatto, S. Evolutionary Generation of Multi-Agent Systems. ArXiv 2026, abs/2602.06511. [Google Scholar]
- Zhang, Y.; Xu, T.; Dai, S.; Shao, Z.; Wu, Q.; Wang, H. EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales. CoRR 2026, abs/2605.11136, [2605.11136. [Google Scholar] [CrossRef]
- Zhang, G.; Niu, L.; Fang, J.; Wang, K.; Bai, L.; Wang, X. Multi-agent Architecture Search via Agentic Supernet. CoRR 2025, abs/2502.04180, 2502.04180. [Google Scholar] [CrossRef]
- Ma, B.; Li, H.; Hu, Z.; Gui, X.; Liu, L.; Liu, S. AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models. CoRR 2025, abs/2510.02669, 2510.02669. [Google Scholar] [CrossRef]
- Yao, T.; Li, Z.; Shen, Z. HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent Systems. CoRR 2026, abs/2602.20229, [2602.20229. [Google Scholar] [CrossRef]
- Guo, J.; Xue, X.; Zhang, L.; Xu, W.; Chen, S.; Torr, P.; Ouyang, W.; Bai, L.; Yin, Z. SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks. 2026. [Google Scholar]
- Fang, W.; Yuan, L.; Lan, G.; Han, D.; Brinton, C.G. Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs. CoRR 2026, abs/2605.08686, [2605.08686. [Google Scholar] [CrossRef]
- Ye, R.; Liu, X.; Wu, Q.; Pang, X.; Yin, Z.; Bai, L.; Chen, S. X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs. CoRR 2025, abs/2505.16997, [2505.16997. [Google Scholar] [CrossRef]
- Feng, Y.; Du, J.; Hong, Y.; Wang, Q.; Yu, L. PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning. CoRR 2025, abs/2508.10501, [2508.10501. [Google Scholar] [CrossRef]
- Su, J.; Xia, Y.; Lan, Q.; Song, X.; Chen, C.; Yang, J.; He, L.; Shi, T. Difficulty-Aware Agent Orchestration in LLM-Powered Workflows. CoRR 2025, abs/2509.11079, [2509.11079. [Google Scholar] [CrossRef]
- Zhuge, M.; Wang, W.; Kirsch, L.; Faccio, F.; Khizbullin, D.; Schmidhuber, J. Language Agents as Optimizable Graphs. CoRR 2024, abs/2402.16823, [2402.16823. [Google Scholar] [CrossRef]
- Li, Z.; Xu, S.; Mei, K.; Hua, W.; Rama, B.; Raheja, O.; Wang, H.; Zhu, H.; Zhang, Y. AutoFlow: Automated Workflow Generation for Large Language Model Agents. CoRR 2024, abs/2407.12821, [2407.12821. [Google Scholar] [CrossRef]
- Zhang, J.; Xiang, J.; Yu, Z.; Teng, F.; Chen, X.; Chen, J.; Zhuge, M.; Cheng, X.; Hong, S.; Wang, J.; et al. AFlow: Automating Agentic Workflow Generation. CoRR 2024, abs/2410.10762, [2410.10762. [Google Scholar] [CrossRef]
- Wang, Y.; Yang, L.; Li, G.; Wang, M.; Aragam, B. ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization. CoRR 2025, abs/2502.04306, [2502.04306. [Google Scholar] [CrossRef]
- Niu, B.; Song, Y.; Lian, K.; Shen, Y.; Yao, Y.; Zhang, K.; Liu, T. Flow: A Modular Approach to Automated Agentic Workflow Generation. CoRR 2025, abs/2501.07834, 2501.07834. [Google Scholar] [CrossRef]
- Zheng, C.; Chen, J.; Lyu, Y.; Ng, W.Z.T.; Zhang, H.; Ong, Y.; Tsang, I.W.; Yin, H. MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming. CoRR 2025, abs/2505.22967, [2505.22967. [Google Scholar] [CrossRef]
- Wang, Y.; Liu, S.; Fang, J.; Meng, Z. EvoAgentX: An Automated Framework for Evolving Agentic Workflows. CoRR 2025, abs/2507.03616, [2507.03616. [Google Scholar] [CrossRef]
- Zhang, G.; Chen, K.; Wan, G.; Chang, H.; Cheng, H.; Wang, K.; Hu, S.; Bai, L. EvoFlow: Evolving Diverse Agentic Workflows On The Fly. CoRR 2025, abs/2502.07373, [2502.07373. [Google Scholar] [CrossRef]
- Liu, S.; Fang, J.; Zhou, H.; Wang, Y.; Meng, Z. SEW: Self-Evolving Agentic Workflows for Automated Code Generation. CoRR 2025, abs/2505.18646, [2505.18646. [Google Scholar] [CrossRef]
- Wei, Y.; Huang, Z.; Li, H.; Xing, W.W.; Lin, T.; He, L. VFlow: Discovering Optimal Agentic Workflows for Verilog Generation. CoRR 2025, abs/2504.03723, 2504.03723. [Google Scholar] [CrossRef]
- Hou, Z.; Tang, J.; Wang, Y. HALO: Hierarchical Autonomous Logic-Oriented Orchestration for Multi-Agent LLM Systems. CoRR 2025, abs/2505.13516, [2505.13516. [Google Scholar] [CrossRef]
- Xu, B.; Ye, Y.; Shen, C.; Zhou, Y.; Chen, C.; Chen, M. HyEvo: Self-Evolving Hybrid Agentic Workflows for Efficient Reasoning. CoRR 2026, abs/2603.19639, [2603.19639. [Google Scholar] [CrossRef]
- Zhu, R.; Jiang, B.; Mei, L.; Yang, F.; Wang, L.; Gao, H.; Bai, F.; Zhao, P.; Lin, Q.; Rajmohan, S.; et al. AdaptFlow: Adaptive Workflow Optimization via Meta-Learning. CoRR 2025, abs/2508.08053, [2508.08053. [Google Scholar] [CrossRef]
- Wang, J.; Xu, S.; Liu, H.; Wang, J.; Luo, Y.; Di, S.; Zhang, M.; Chen, L. Learning to Compose for Cross-domain Agentic Workflow Generation. CoRR 2026, abs/2602.11114, [2602.11114. [Google Scholar] [CrossRef]
- Yuan, B.; Zhou, Y.; Xu, Z.; Ramnath, K.; Feng, A.; Srinivasan, B. BayesFlow: A Probability Inference Framework for Meta-Agent Assisted Workflow Generation. CoRR 2026, abs/2601.22305, [2601.22305. [Google Scholar] [CrossRef]
- Kong, M.; Qu, Z.; Zhou, Z.; Liang, P.; Li, X.; Shang, Z.; Hong, Z.; Huang, K.; Wang, Z.; Dai, Z. Workflow-R1: Group Sub-sequence Policy Optimization for Multi-turn Workflow Construction. CoRR 2026, abs/2602.01202, 2602.01202. [Google Scholar] [CrossRef]
- Ma, Z.; Zhao, Z.; Hua, C.; Berto, F.; Park, J. JudgeFlow: Agentic Workflow Optimization via Block Judge. abs/2601.07477; CoRR. 2026; p. 2601.07477. [Google Scholar] [CrossRef]
- Xu, S.; Zhang, J.; Di, S.; Luo, Y.; Yao, L.; Liu, H.; Zhu, J.; Liu, F.; Zhang, M. RobustFlow: Towards Robust Agentic Workflow Generation. CoRR 2025, abs/2509.21834, [2509.21834. [Google Scholar] [CrossRef]
- Shi, X.; Zheng, M.; Lou, Q. Learning Latency-Aware Orchestration for Parallel Multi-Agent Systems. CoRR 2026, abs/2601.10560, [2601.10560. [Google Scholar] [CrossRef]
- Li, J.; Hong, Z.; Shen, M.; Zhang, Y.; Gan, C. FlowCompile: An Optimizing Compiler for Structured LLM Workflows. CoRR 2026, abs/2605.13647, [2605.13647. [Google Scholar] [CrossRef]
- Li, A.; Yang, S.; Chen, F.; Xu, T.; Li, P.; Su, Z. GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving. CoRR 2026, abs/2605.22566, [2605.22566. [Google Scholar] [CrossRef]
- Li, J.; Zhang, E.; Zhou, D.; Chen, E.; Yan, Y. Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints. CoRR 2026, abs/2605.19140, [2605.19140. [Google Scholar] [CrossRef]
- Liu, S.; Li, M.; Fu, D.; Wang, H.P.; Xia, Y.; Li, H.; Yan, H.; Li, P. Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows. 2026. [Google Scholar]
- Hu, Y.; Cai, Y.; Du, Y.; Zhu, X.; Liu, X.; Yu, Z.; Hou, Y.; Tang, S.; Chen, S. Self-Evolving Multi-Agent Collaboration Networks for Software Development. CoRR 2024, abs/2410.16946, [2410.16946. [Google Scholar] [CrossRef]
- Zhang, G.; Yue, Y.; Sun, X.; Wan, G.; Yu, M.; Fang, J.; Wang, K.; Cheng, D. G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks. CoRR 2024, abs/2410.11782, [2410.11782. [Google Scholar] [CrossRef]
- Zhang, G.; Yue, Y.; Li, Z.; Yun, S.; Wan, G.; Wang, K.; Cheng, D.; Yu, J.X.; Chen, T. Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems. CoRR 2024, abs/2410.02506, [2410.02506. [Google Scholar] [CrossRef]
- Li, B.; Zhao, Z.; Lee, D.; Wang, G. Adaptive Graph Pruning for Multi-Agent Communication. CoRR 2025, abs/2506.02951, 2506.02951. [Google Scholar] [CrossRef]
- Zhang, R.; Zhao, X.; Wang, R.; Chen, S.; Zhang, G.; Zhang, A.; Wang, K.; Wen, Q. SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communication. CoRR 2025, abs/2508.11733, [2508.11733. [Google Scholar] [CrossRef]
- Leong, H.Y.; Li, Y.; Wu, Y.; Ouyang, W.; Zhu, W.; Gao, J.; Han, W. AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System. CoRR 2025, abs/2510.01617, 2510.01617. [Google Scholar] [CrossRef]
- Jiang, E.H.; Wan, G.; Yin, S.; Li, M.; Wu, Y.; Liang, X.; Li, X.; Sun, Y.; Wang, W.; Chang, K.; et al. Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models. CoRR 2025, abs/2510.07799, [2510.07799. [Google Scholar] [CrossRef]
- Li, S.; Liu, Y.; Zheng, Y.; Li, M.; Nguyen, Q.V.H.; Pan, S. OFA-MAS: One-for-All Multi-Agent System Topology Design based on Mixture-of-Experts Graph Generative Models. CoRR 2026, abs/2601.12996, [2601.12996. [Google Scholar] [CrossRef]
- Wu, X.; Liu, X.; Lu, J.; Wang, S.; Qiu, X.; Shu, Y.; Hu, J.; Guo, C.; Yang, B. ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies. CoRR 2026, abs/2602.14681, 2602.14681. [Google Scholar] [CrossRef]
- Jiang, E.H.; Li, L.; Sun, R.; Liang, X.; Li, Y.; Wu, Y.; Luo, H.; Li, H.; Zhang, Z.; Kang, Z.; et al. Agent Q-Mix: Selecting the Right Action for LLM Multi-Agent Systems through Reinforcement Learning. CoRR 2026, abs/2604.00344, 2604.00344. [Google Scholar] [CrossRef]
- Zhang, Z.; Zhou, W.; Li, J.; Fei, H.; Wen, J.; Ji, W. RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation. CoRR 2026, abs/2605.09907, [2605.09907. [Google Scholar] [CrossRef]
- Wu, X.; Lu, J.; Yan, S.; Qiu, X.; Hu, J.; Guo, C.; Yang, B. Differentiable Mixture-of-Agents Incentivizes Swarm Intelligence of Large Language Models. CoRR 2026, abs/2605.15706, [2605.15706. [Google Scholar] [CrossRef]
- Wang, S.; Lu, R.; Yang, Z.; Wang, Y.; Zhang, Y.; Xu, L.; Xu, Q.; Yin, G.; Chen, C.; Guan, X. AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation. CoRR 2026, abs/2602.17100, [2602.17100. [Google Scholar] [CrossRef]
- Xu, C.; Hu, Y.; Wang, R.; Lin, X.; Wang, W.; Liu, D.; Feng, F. TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems. CoRR 2026, abs/2605.09539, [2605.09539. [Google Scholar] [CrossRef]
- Zhang, T.; Zhou, Z.; Wan, J.; Hu, T.; Wang, C.; He, X.; Hong, R. Learning Transferable Topology Priors for Multi-Agent LLM Collaboration Across Domains. CoRR 2026, abs/2605.17359, [2605.17359. [Google Scholar] [CrossRef]
- Wang, X.; Wang, J.; Zhang, F.; Hu, Y.; Zhang, D.; Ye, Y.; Ban, Y.; Han, J.; Wang, R. MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer. CoRR 2026, abs/2605.17361, [2605.17361. [Google Scholar] [CrossRef]
- Gou, W.; Liu, Z. Dynamic Trust-Aware Sparse Communication Topology for LLM-Based Multi-Agent Consensus. 2026. [Google Scholar]
- Talluri, A.; Anne, P.; Pendiyala, B.C.; Chilukuri, R. Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation. CoRR 2026, abs/2605.05657, [2605.05657. [Google Scholar] [CrossRef]
- Tastan, N.; Iacob, A.; Sani, L.; Kurmanji, M.; Lane, N.D.; Horvath, S.; Nandakumar, K. Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems. CoRR 2026, abs/2605.15573, [2605.15573. [Google Scholar] [CrossRef]
- Wang, D.; Yin, D.; Desai, R.; Li, L.; Celikyilmaz, A.; Ni, A. Learning to Interrupt in Language-based Multi-agent Communication. CoRR 2026, abs/2604.06452, [2604.06452. [Google Scholar] [CrossRef]
- Yu, Y.; Liu, H.; Jin, H.; Yuan, X.; Kuang, P.; Wang, H. Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems. CoRR 2026, abs/2604.21794. [Google Scholar] [CrossRef]
- Xu, T.; Wen, H.; Li, M. Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents. CoRR 2026, abs/2605.22166, [2605.22166. [Google Scholar] [CrossRef]
- Chen, M.; Wang, J.; Liu, Z.; Wang, Y.; Wang, Q. From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws. 2026. [Google Scholar] [CrossRef]
- Wei, C.; Gao, M.; Han, Z.; Chen, K.; Zhuang, Y.; Guan, H.; Zhang, Y.; Cheng, Y.; He, J.; Chen, H.; et al. The World Leaks the Future: Harness Evolution for Future Prediction Agents. CoRR 2026, abs/2604.15719, [2604.15719. [Google Scholar] [CrossRef]
- Zhang, H.; Zhang, S.; Li, K.; Zhang, C.; Chen, Y.; Zhang, Y.; Bai, L.; Hu, S. Self-Harness: Harnesses That Improve Themselves. 2026. [Google Scholar] [CrossRef]
- Kakade, A.; Srivastava, V.; Karande, S.S. Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair. CoRR 2026, abs/2603.23129, [2603.23129. [Google Scholar] [CrossRef]
- Lee, Y.; Nair, R.; Zhang, Q.; Lee, K.; Khattab, O.; Finn, C. Meta-Harness: End-to-End Optimization of Model Harnesses. CoRR 2026, abs/2603.28052, [2603.28052. [Google Scholar] [CrossRef]
- Liu, H.; Shou, C.; Liu, X.; Wen, H.; Chen, Y.; Fang, R.J.; Feng, Y. Synthesizing Multi-Agent Harnesses for Vulnerability Discovery. CoRR 2026, abs/2604.20801. [Google Scholar] [CrossRef]
- Zhang, D. AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering. CoRR 2026, abs/2601.04620, [2601.04620. [Google Scholar] [CrossRef]
- Nanda, R.; Maddila, C.; Jha, S.; Khan, E.M.; Paltenghi, M.; Chandra, S. Wink: Recovering from Misbehaviors in Coding Agents. CoRR 2026, abs/2602.17037, [2602.17037. [Google Scholar] [CrossRef]
- Mulian, H.; Zeltyn, S.; Levy, I.; Galanti, L.; Yaeli, A.; Shlomov, S. AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems. CoRR 2026, abs/2603.29848, [2603.29848. [Google Scholar] [CrossRef]
- Bonagiri, A.; Borkar, D.; Anderias, G.J.; Rafatirad, S.; Homayoun, H. CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures. CoRR 2026, abs/2605.25338, [2605.25338. [Google Scholar] [CrossRef]
- Zhang, G.; Wang, J.; Chen, J.; Zhou, W.; Wang, K.; Yan, S. AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems? CoRR 2025, abs/2509.03312, [2509.03312. [Google Scholar] [CrossRef]
- Zhang, S.; Yin, M.; Zhang, J.; Liu, J.; Han, Z.; Zhang, J.; Li, B.; Wang, C.; Wang, H.; Chen, Y.; et al. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems. CoRR 2025, abs/2505.00212, [2505.00212. [Google Scholar] [CrossRef]
- Zelikman, E.; Lorch, E.; Mackey, L.; Kalai, A.T. Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. CoRR 2023, abs/2310.02304, 2310.02304. [Google Scholar] [CrossRef]
- Yin, X.; Wang, X.; Pan, L.; Lin, L.; Wan, X.; Wang, W.Y. Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025. [Google Scholar]
- Robeyns, M.; Szummer, M.; Aitchison, L. A Self-Improving Coding Agent. CoRR 2025, abs/2504.15228, [2504.15228. [Google Scholar] [CrossRef]
- Cai, Q.; Zhang, Y.; Jia, X.; Zheng, H.; Xue, W.; Song, J.; Tian, X.; Guo, Y. MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems. CoRR 2026, abs/2605.22794, [2605.22794. [Google Scholar] [CrossRef]
- Robol, M.; Giorgini, P. Self-Evolving Software Agents. CoRR 2026, abs/2604.27264, [2604.27264. [Google Scholar] [CrossRef]
- Che, L.; Yang, Y.; Lin, P.; Wang, C.; Wang, X.; Su, J. DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations. CoRR 2026, abs/2605.24539, [2605.24539. [Google Scholar] [CrossRef]
- Ho, M.; Liu, B.; Chen, J.; Wang, A.X.; Qin, L. SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation. 2026. [Google Scholar]
- Wang, R.; Huang, J.; Wang, P.; Liu, X.; Kong, L.; Zhang, T. Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory. 2026. [Google Scholar]
- Zhang, S.; Yuan, C.; Guo, R.; Yu, X.; Xu, R.; Chen, Z.; Li, Z.; Yang, Z.; Guan, S.; Tang, Z.; et al. EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines. CoRR 2026, abs/2601.09465, [2601.09465. [Google Scholar] [CrossRef]
- Zhang, J.; Hu, S.; Lu, C.; Lange, R.T.; Clune, J. Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Wang, W.; Piekos, P.; Li, N.; Laakom, F.; Chen, Y.; Ostaszewski, M.; Zhuge, M.; Schmidhuber, J. Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine. CoRR 2025, abs/2510.21614, [2510.21614. [Google Scholar] [CrossRef]
- Weng, Z.; Antoniades, A.; Nathani, D.; Zhang, Z.; Pu, X.; Wang, X.E. Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing. CoRR 2026, abs/2602.04837, [2602.04837. [Google Scholar] [CrossRef]
- Zhang, J.; Gu, Y.; Ruan, J.; Song, M.; Peng, Y.; Han, Z.; Xiang, J.; Wang, Z.; Yang, C.; Ouyang, Y.; et al. Harnessing Agentic Evolution. CoRR 2026, abs/2605.13821, [2605.13821. [Google Scholar] [CrossRef]
- Qu, Y.; Lu, M. Bilevel Autoresearch: Meta-Autoresearching Itself. CoRR 2026, abs/2603.23420, [2603.23420. [Google Scholar] [CrossRef]
- Liu, Z.; Shi, Z.; Sang, Y.; He, B.; Lin, M.; Wei, T.; Wang, D.; Dumoulin, B.; Jin, W.; Lu, H. Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams. 2026. [Google Scholar]
- Liu, S.; Agarwal, S.; Maheswaran, M.; Cemri, M.; Li, Z.; Mang, Q.; Naren, A.; Boneh, E.; Cheng, A.; Pan, M.Z.; et al. EvoX: Meta-Evolution for Automated Discovery. CoRR 2026, abs/2602.23413, [2602.23413. [Google Scholar] [CrossRef]
- Wu, X.; Yin, S.; Kang, Y.; Zhang, X.; Xu, Q.; Chen, Z.; Zhang, W. SGM: A Statistical Godel Machine for Risk-Controlled Recursive Self-Modification. CoRR 2025, abs/2510.10232, [2510.10232. [Google Scholar] [CrossRef]
- Garralda-Barrio, M. Governed Evolution of Agent Runtimes through Executable Operational Cognition. CoRR 2026, abs/2605.27328, [2605.27328. [Google Scholar] [CrossRef]
- Hakim, S.B.; Guo, K.; Tan, W.; Velasquez, A.; Xu, S.; Song, H.H. ANNEAL: Adapting LLM Agents via Governed Symbolic Patch Learning. CoRR 2026, abs/2605.16309, [2605.16309. [Google Scholar] [CrossRef]
- Sahoo, S.; Chadha, A.; Jain, V.; Chaudhary, D. SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement. CoRR 2026, abs/2603.06333, [2603.06333. [Google Scholar] [CrossRef]
- Shi, D.; He, J.; Chen, J.; Wang, B.; Nakashima, Y. Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems. 2026. [Google Scholar]
- Zala, A.; Cho, J.; Lin, H.; Yoon, J.; Bansal, M. EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents. arXiv 2024, arXiv:cs. [Google Scholar]
- Yeo, T.; Weerakoon, D.; Weerakoon, D.; Misra, A. Towards Adaptive Environment Generation for Training Embodied Agents, 2026. arXiv arXiv:cs.
- Parker-Holder, J.; Jiang, M.; Dennis, M.; Samvelyan, M.; Foerster, J.; Grefenstette, E.; Rocktäschel, T. Evolving Curricula with Regret-Based Environment Design, 2022. arXiv arXiv:cs.
- Garcin, S.; Doran, J.; Guo, S.; Lucas, C.G.; Albrecht, S.V. DRED: Zero-Shot Transfer in Reinforcement Learning via Data-Regularised Environment Design. arXiv 2024, arXiv:cs. [Google Scholar]
- Kang, H.; Ye, X.; Liu, Y.; Mantri, S.H.; Mao, L.; Fleming, J.; Regmi, D.; Qin, L. SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Chen, Z.; Zhao, Z.; Zhang, K.; Liu, B.; Qi, Q.; Wu, Y.; Kalluri, T.; Cao, S.; Xiong, Y.; Tong, H.; et al. Scaling Agent Learning via Experience Synthesis. arXiv 2025, arXiv:cs. [Google Scholar]
- Qi, Z.; Liu, X.; Iong, I.L.; Lai, H.; Sun, X.; Zhao, W.; Yang, Y.; Yang, X.; Sun, J.; Yao, S.; et al. WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning. arXiv 2024, arXiv:cs. [Google Scholar]
- Yang, S.; Ma, Z.; Huang, T.; Hu, Y.; Wang, Y.; Chu, X. CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, San Diego, California, United States, 2026; Volume 1, pp. 23015–23036. [Google Scholar] [CrossRef]
- Fang, W.; Liu, S.; Zhou, Y.; Zhang, K.; Zheng, T.; Chen, K.; Song, M.; Tao, D. SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data, 2026. arXiv arXiv:cs.
- Huang, J.; Li, Z.; Hu, Y.; Zhang, Z.; Coates, M.; Quan, X.; Zhang, Y. Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation. arXiv 2026, arXiv:cs. [Google Scholar]
- Amjith, S.; Wang, M.X.; Lynch, J.; Gundlach, H.; Thompson, N. SAGE: Self-play Adversarial Games Enhance Large Language Model Reasoning Capabilities. [PubMed]
- Chen, Z.; Deng, Y.; Yuan, H.; Ji, K.; Gu, Q. Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models. arXiv 2024, arXiv:cs. [Google Scholar]
- Luo, H.; Sun, Q.; Xu, C.; Zhao, P.; Lin, Q.; Lou, J.; Chen, S.; Tang, Y.; Chen, W. Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena. arXiv 2024, arXiv:cs. [Google Scholar]
- Zeng, Y.; Cui, X.; Jin, X.; Mi, Q.; Liu, G.; Sun, Z.; Yang, M.; Li, D.; Ma, W.; Yang, N.; et al. Evolving LLMs’ Self-Refinement Capability via Synergistic Training-Inference Optimization. arXiv 2025, arXiv:cs. [Google Scholar]
- Qu, Y.; Zhang, T.; Garg, N.; Kumar, A. Recursive Introspection: Teaching Language Model Agents How to Self-Improve. arXiv 2024, arXiv:cs. [Google Scholar]
- Yang, H.; Le, K.; Hua, T.; Gao, S.; Xu, B.; Tang, Z.; Xu, J.; Chawla, N.V.; Jin, H.; Srinivasan, V. Dynamic Noise Preference Optimization: Self-Improvement of Large Language Models with Self-Synthetic Data, 2026. arXiv arXiv:cs.
- Xiao, W.; Lin, H.; Peng, A.; Xue, H.; He, T.; Xie, Y.; Hu, F.; Wu, J.; Luo, Z.; Fan, L.J.; et al. Self-Improving Vision-Language-Action Models with Data Generation via Residual RL. arXiv 2025, arXiv:cs. [Google Scholar]
- Lin, I.W.; Hu, Y.; Li, S.S.; Geng, S.; Koh, P.W.; Zettlemoyer, L.; Althoff, T.; Ghazvininejad, M. Self-Improving VLM Judges Without Human Annotations. arXiv 2025, arXiv:cs. [Google Scholar]
- Tang, C.; Huang, H.Y.; Liu, W.; Zheng, J.; Yang, S.; Wu, Y. Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model. arXiv 2026, arXiv:2604.17739. [Google Scholar]
- Cheng, Y.; Wang, Z.; Ma, W.; Zhu, W.; Deng, Y.; Zhao, J. EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, Q.; Ruan, S.; Upasani, S.; Hong, F.; Ji, C.; Hu, C.; Li, B.; Li, H.; Olukotun, K. Learning What to Learn: Curriculum Curation for Test-Time Agent Learning. In Proceedings of the ICLR 2026 Workshop on AI with Recursive Self-Improvement (RSI), 2026. [Google Scholar]
- Bukkapatnam, K.; Lala, A.; Patel, L. Adaptive Meta-Curriculum for Test-Time Self-Improvement. In Proceedings of the ICLR 2026 Workshop on AI with Recursive Self-Improvement (RSI), 2026. [Google Scholar]
- Gu, Z.; Light, J.; Astudillo, R.; Ye, Z.; He, L.; Zou, H.P.; Cheng, W.; Paternain, S.; Yu, P.S.; Yue, Y. Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training, 2026. arXiv arXiv:cs.
- Liu, J.; Xiong, K.; Xia, P.; Zhou, Y.; Ji, H.; Feng, L.; Han, S.; Ding, M.; Yao, H. Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Huang, Y.; Yu, X.; Wei, Z. ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization, 2026. arXiv arXiv:cs.
- Dong, G.; Lu, J.; Huang, J.; Zhong, W.; Liu, L.; Huang, S.; Li, Z.; Zhao, Y.; Song, X.; Li, X.; et al. Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence. arXiv 2026, arXiv:cs. [Google Scholar]
- Xia, P.; Zeng, K.; Liu, J.; Qin, C.; Wu, F.; Zhou, Y.; Xiong, C.; Yao, H. Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning. arXiv 2025, arXiv:cs. [Google Scholar]
- Huang, C.; Yu, W.; Wang, X.; Zhang, H.; Li, Z.; Li, R.; Huang, J.; Mi, H.; Yu, D. R-Zero: Self-Evolving Reasoning LLM from Zero Data, 2025. arXiv arXiv:cs.
- Sundaram, S.; Quan, J.; Kwiatkowski, A.; Ahuja, K.; Ollivier, Y.; Kempe, J. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability, 2026. arXiv arXiv:cs.
- Jana, S.; Sancaktar, C.; Daniš, T.; Martius, G.; Orvieto, A.; Kolev, P. GASP: Guided Asymmetric Self-Play For Coding LLMs, 2026. arXiv arXiv:cs.
- Shao, R.; Asai, A.; Shen, S.Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S.G.; Sontag, D.; et al. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, G.; Mishra, B.D.; Wang, Z.; Yan, J.; Chen, Y.; Li, C.L.; Le, L.T.; Han, R.; Lee, G.; Tong, H.; et al. RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards. arXiv 2026, arXiv:cs. [Google Scholar]
- Liu, Z.; Zhang, L.; Wang, X.; Xu, Z.; Zhan, S.; Shan, X.; Huang, W.; Dai, T.; Xia, S.T.; Huo, C.; et al. ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents. arXiv 2026, arXiv:cs. [Google Scholar]
- Zheng, C.; Mo, X.; Ma, X.; Lin, Q.; Zhao, Y.; Zhu, J.; Lou, X.; Wang, J.; Wang, Z.; Liu, W.; et al. Adaptive Milestone Reward for GUI Agents. arXiv 2026, arXiv:2602.11524. [Google Scholar]
- Kim, Z.M.; Park, C.; Raheja, V.; Kim, S.; Kang, D. Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models. arXiv 2025, arXiv:2504.20157. [Google Scholar]
- Ding, H.; Huang, B.; Fang, Y.; Liao, W.; Li, Z.; Zhang, J.; Wu, Z.; Zhao, J.; Wang, Y. EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning, 2026. arXiv arXiv:cs.
- Li, S.S.; Xin, R.; Xiao, T.; Wang, Y.; Shao, R.; Hao, Z.; Sclar, M.; Oh, S.; Brahman, F.; Koh, P.W.; et al. EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics, 2026. arXiv arXiv:cs.
- Sheng, L.; Ma, W.; Hong, R.; Wang, X.; Zhang, A.; Chua, T.S. Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics. arXiv 2026, arXiv:cs. [Google Scholar]
- Guan, X.; Hu, X.; Huang, S.; Wang, Z.; Zhang, B.; Li, Z.; Xie, P.; Liu, B.; Cao, J. EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation, 2026. arXiv arXiv:cs.
- Tian, Z.; Zhang, J.; Li, R.; Bo, X.; Li, Y.; Chen, X. ARCO: Adaptive Rubric with Co-Evolution for Multi-Step LLM-Based Agents. arXiv 2026, arXiv:cs. [Google Scholar]
- Feng, A.Z.; Wang, C.; Wen, B.; Wang, Y.; Luo, Y.; Wang, H.; Huang, M. RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models, 2026. arXiv arXiv:cs.
- Wang, X.; Wu, T.; Tang, M.; Li, J.; Liu, Q.; Zheng, Z. The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement, 2026. arXiv arXiv:cs.
- Shi, T.; Huang, C.; Wan, F.; Zhong, L.; Yang, Z.; Shen, W.; Quan, X.; Yan, M. Mutual-Taught for Co-adapting Policy and Reward Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, Z.; Cao, Z.; Huang, W.; Zhang, Y.; Qi, K.; Wang, R.; Zheng, Z.; Zhao, J.; Zhu, H.; Wu, H.; et al. MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback Reflux, 2026. arXiv arXiv:cs.
- Lin, Y.; Wang, L.; Lin, K.; Lin, Z.; Gong, K.; Li, W.; Lin, B.; Li, Z.; Zhang, S.; Peng, Y.; et al. JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization. arXiv 2025, arXiv:cs. [Google Scholar]
- Li, Z.; Jiang, L.; Hu, Y.; Zeng, X.; Li, Y.; Zhang, X.; Chen, G.; Pan, Z.; Li, X.; Liu, Y. No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning, 2026. arXiv arXiv:cs.
- Lin, J.; Yu, X.; Xin, Y.; Guo, Y.; Jiang, Z.; Yue, Z.; Wang, W.; Zou, H.; Qin, C.; Xiong, H. ICRL: Learning to Internalize Self-Critique with Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Wu, M.; Zhang, G.; Min, S.; Levine, S.; Kumar, A. RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks. arXiv 2025, arXiv:2511.01758. [Google Scholar]
- Sun, W.; Cheng, X.; Fan, J.; Yu, X.; Xu, Y.; He, S.; Zhao, J.; Liu, K. Towards Agentic Self-Learning LLMs in Search Environment. arXiv 2025, arXiv:cs. [Google Scholar]
- Sui, Y.; Hooi, B. Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation. arXiv 2026, arXiv:cs. [Google Scholar]
- Pan, T.; Yan, Y.; Wang, Z.; Zhang, R.; Hou, G.; Zhang, W.; Lu, W.; Xiao, J.; Shen, Y. CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution. arXiv 2026, arXiv:2603.17775. [Google Scholar]
- Zhu, H.; Cai, C.; Song, Y.; Chen, X.; Han, S.; Guo, Y. Self-Evolving Deep Research via Joint Generation and Evaluation. arXiv 2026, arXiv:2606.04507. [Google Scholar]
- Lu, S.; Wang, H.; Chen, Z.; Tang, Y. URPO: A Unified Reward & Policy Optimization Framework for Large Language Models. arXiv 2025, arXiv:2507.17515. [Google Scholar]
- Zha, K.; Gao, Z.; Shen, M.; Hong, Z.W.; Boning, D.S.; Katabi, D. RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning. arXiv 2025, arXiv:2505.15034. [Google Scholar]
- Wang, Y.; Xie, T.; Shen, K.; Wang, M.; Yang, L. RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System, 2026. arXiv arXiv:cs.
- Zhang, Y.; Fang, M.; Chen, Z.; Pechenizkiy, M. Self-evolving LLM agents with in-distribution Optimization, 2026. arXiv arXiv:cs.
- Guan, X.; Zhang, L.L.; Liu, Y.; Shang, N.; Sun, Y.; Zhu, Y.; Yang, F.; Yang, M. rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking. arXiv 2025, arXiv:2501.04519. [Google Scholar]
- Zhou, C.; Xu, T.; Lin, J.; Ge, D. StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models. arXiv 2025, arXiv:2509.22558. [Google Scholar]
- Zhang, J.; Ma, G.; Liu, S.; Hu, Z.; Jing, Y.; Lin, T.E.; Li, Y.; Tao, D. STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning. arXiv 2026, arXiv:2605.18851. [Google Scholar]
- Xiao, H.; Wang, G.; Chai, Y.; Lu, Z.; Lin, W.; He, H.; Fan, L.; Bian, L.; Hu, R.; Liu, L.; et al. UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhi, X.; zhou, P.; Lu, C.; Lv, H.; Liang, Y.; Zhang, R.; Gao, Y.; WU, Y.; Hu, Y.; Gu, H.; et al. SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility. arXiv 2026, arXiv:2604.07837. [Google Scholar]
- Fan, L.; Chen, M.; Zhu, T.; Liu, K.; Xia, X.; Li, S.; Liu, Z. ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? arXiv 2026, arXiv:2604.07864. [Google Scholar]
- Wang, H.; Wang, G.; Xiao, H.; Zhou, Y.; Pan, Y.; Wang, J.; Xu, K.; Wen, Y.; Ruan, X.; Chen, X.; et al. Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents, 2026. arXiv arXiv:cs.
- Tu, S.; Xu, C.; Zhang, Q.; Ma, Y.; Zhang, Y.; Li, L.; Li, D.; Lan, X.; Zhao, D. UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation. arXiv 2026, arXiv:cs. [Google Scholar]
- Jeon, U.; Kwon, J.; Sullivan, M.A.; Lee, C.E.; Lin, G. ATLAS: Adaptive Self-Evolutionary Research Agent with Task-Distributed Multi-LLM Supporters. arXiv 2026, arXiv:cs. [Google Scholar]
- Rao, A.; Advani, N.K. AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes. arXiv 2026, arXiv:cs. [Google Scholar]
- Dong, H.; Yang, D.; Liang, X.; Feng, C.; Ran, J. AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining. Proc. Adv. Neural Inf. Process. Syst. 2025, arXiv:cs. [Google Scholar] [CrossRef]
- Lin, Z.; Xue, C.; Liang, D.; Han, X.; Liu, P.; Wu, X.; Jiang, L.; Lu, Y.; Shi, H.; Liang, S.; et al. Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning. arXiv 2026, arXiv:2604.14010. [Google Scholar]
- Sakip, A.; Fuadi, E.H.; Sayedelahl, O.; Li, Z.; She, J.; Aji, A.F.; Liu, S.; Xing, E.; Ho, Q. COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training, 2026. arXiv arXiv:cs.
- Mai, L.; Li, G.; Wagenländer, M.; Fertakis, K.; Brabete, A.O.; Pietzuch, P. KungFu: Making Training in Distributed Machine Learning Adaptive. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), 2020; USENIX Association; pp. 937–954. [Google Scholar]
- Qiao, A.; Choe, S.K.; Subramanya, S.J.; Neiswanger, W.; Ho, Q.; Zhang, H.; Ganger, G.R.; Xing, E.P. Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In Proceedings of the 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21), 2021; USENIX Association; pp. 1–18. [Google Scholar]
- Subramanya, S.J.; Arfeen, D.; Lin, S.; Qiao, A.; Jia, Z.; Ganger, G.R. Sia: Heterogeneity-aware, Goodput-optimized ML-cluster Scheduling. In Proceedings of the Proceedings of the 29th Symposium on Operating Systems Principles, 2023; pp. 642–657. [Google Scholar] [CrossRef]
- Bian, Z.; Li, S.; Wang, W.; You, Y. Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters. In Proceedings of the Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2021; pp. 1–15, [2108.03645. [Google Scholar] [CrossRef]
- Dai, Y.; He, K.; Wang, A. DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, L.; Jia, T.; Zhai, Y.; Fang, L.; Zheng, K.; Liu, H.; Huang, X.; Yu, P.S.; Li, Y. Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning. arXiv 2026, arXiv:2605.04431. [Google Scholar]
- Tan, Z.; Abdullahi, M.; Shi, T.; Yuan, H.; Xu, Z.; Yu, C.; Li, B.; Zhao, B. EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models. arXiv 2025, arXiv:cs. [Google Scholar]
- Peng, Y.; Bao, Y.; Chen, Y.; Wu, C.; Guo, C. Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters. In Proceedings of the Proceedings of the Thirteenth EuroSys Conference, 2018. [Google Scholar] [CrossRef]
- Zhang, X.; Zhao, H.; Xiao, W.; Jia, X.; Xu, F.; Li, Y.; Lin, W.; Liu, F. Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling. arXiv 2024, arXiv:cs. [Google Scholar]
- Fang, H.; Zhu, W.; Han, B.; Zhang, A.; Pan, Z.; Yang, S.; Zhang, S.; Gai, J.; Tang, P.; Hu, C.; et al. LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents. arXiv 2026, arXiv:2606.18388. [Google Scholar]
- Chen, G.; Shi, Y.; Li, Y.; Li, B.; Xu, X.; Wei, H.; Ni, S.; Yang, M.; Ye, J. EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning, 2026. arXiv arXiv:cs.
- Yu, Z.; Yin, P.; Gao, S.; He, S.; Cai, K.; Zhang, X.P. AutoTrainess: Teaching Language Models to Improve Language Models Autonomously. arXiv 2026, arXiv:2606.31551. [Google Scholar]
- Li, Q.; Zhang, Y.; Yang, X.; Yang, X.; Wang, Z.; Liu, W.; Bian, J. FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents. arXiv 2026, arXiv:2603.01712. [Google Scholar]
- Ning, J.; Li, X.; Zeng, J.; Kang, H.; Xiong, C. Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes. arXiv 2026, arXiv:cs. [Google Scholar]
- Mroueh, Y.; Fonseca, C.; Belgodere, B.; Cox, D. CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery, 2026. arXiv arXiv:cs.
- Chen, S.; Tang, Z.; Zhang, W.; Yang, F.; Wang, Y.; Li, T.; Liu, Y. OptiCo: Adaptive Distributed Training Optimization via Collaborative Agent Reasoning. In Proceedings of the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 2026. [Google Scholar]
- Gao, S.; Fang, A.; Zitnik, M. AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation. arXiv 2026, arXiv:2605.28655. [Google Scholar]
- Ma, Z.; Wang, G.; Xie, X.; Chen, Y.; Du, H.; Li, B.; Sun, Y.; Liu, W.; Chen, K.; Li, Y. TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration, 2026. arXiv arXiv:cs.
- Xia, S.; Zhang, Y.; Chen, A.; Wu, S.; Yuan, S.; Xiao, Y. From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents, 2026. arXiv arXiv:cs.
- Ma, Y.J.; Liang, W.; Wang, G.; Huang, D.A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, L.J.; Anandkumar, A. Eureka: Human-Level Reward Design via Coding Large Language Models. arXiv 2023, arXiv:cs. [Google Scholar]
- Sun, S.; Liu, R.; Lyu, J.; Yang, J.W.; Zhang, L.; Li, X. A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning. arXiv 2024, arXiv:2410.14660. [Google Scholar]
- Wu, X.; Zheng, Z.; Xiong, H. LLM-ALSO: LLM-Driven Adaptive Learning-Signal Optimization for Multi-Agent Reinforcement Learning. arXiv 2026, arXiv:2605.29293. [Google Scholar]
- Hazra, R.; Sygkounas, A.; Persson, A.; Loutfi, A.; Martires, P.Z.D. REvolve: Reward Evolution with Large Language Models using Human Feedback. arXiv 2024, arXiv:2406.01309. [Google Scholar]
- Du, S.; Yan, X.; Shi, J.; Cao, Z.; Feng, S.; Liang, Z.; Sun, B.; Peng, T.; Zhou, Y.; Li, X.; et al. MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery. arXiv 2026, arXiv:2606.06473. [Google Scholar]
- Wang, H.; Wu, Y.; Chang, D.; Wei, L.; Heldt, L. Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents, 2026. arXiv arXiv:cs.
- Li, P.; Tang, H.; Qiao, J.; Zheng, Y.; Hao, J. LaRes: Evolutionary Reinforcement Learning with LLM-based Adaptive Reward Search. In Proceedings of the Advances in Neural Information Processing Systems, 2025. [Google Scholar]
- Li, P.; Hao, J.; Tang, H.; Yuan, Y.; Qiao, J.; Dong, Z.; Zheng, Y. R*: Efficient Reward Design via Reward Structure Evolution and Parameter Alignment Optimization with Large Language Models. In Proceedings of the Proceedings of the 42nd International Conference on Machine Learning. PMLR, Proceedings of Machine Learning Research. 2025; Vol. 267, pp. 34509–34527. [Google Scholar]
- Gao, N.; Zhang, X.; Jiang, X.; You, M.; Zhang, M.; Deng, Y. RF-Agent: Automated Reward Function Design via Language Agent Tree Search. arXiv 2026, arXiv:2602.23876. [Google Scholar]
- Huang, C.; Chang, Y.; Lin, J.; Liang, J.; Zeng, R.; Li, J. Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution. arXiv 2024, arXiv:2412.13492. [Google Scholar]
- Liu, Z.; Chai, J.; Zhu, X.; Tang, S.; Ye, R.; Zhang, B.; Bai, L.; Chen, S. ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering. arXiv 2025, arXiv:cs. [Google Scholar]
- Zhang, Y.; Zhou, K.; Xu, Z.; Ramnath, K.; Zhou, Y.; Woo, S.; Ding, H.; Cheong, L.L. Learning to Ideate for Machine Learning Engineering Agents, 2026. arXiv arXiv:cs.
- Wu, F.; Zheng, X.; Wang, Z.; ming Dai, Y.; Li, H. RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses, 2026. arXiv arXiv:cs.
- Chen, M.; Xiao, B.; Liang, D.; Zeng, C.; Wen, Z. Efficient Hyperparameter Optimization for LLM Reinforcement Learning. arXiv 2026, arXiv:2606.03073. [Google Scholar]
- Lu, C.; Holt, S.; Fanconi, C.; Chan, A.J.; Foerster, J.; van der Schaar, M.; Lange, R.T. Discovering Preference Optimization Algorithms with and for Large Language Models. Proc. Adv. Neural Inf. Process. Syst. 2024, 2406.08414. [Google Scholar]
- Cheng, S.; Li, T.; Huang, X.; Yin, X.; Zou, D. Differentiable Evolutionary Reinforcement Learning. arXiv 2026, arXiv:cs. [Google Scholar]
- Ahmadi, A.; Sharif, S.; Banad, Y. Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning, 2026. arXiv arXiv:cs.
- Han, X.; Yang, Q.; Chen, X.; Chu, X.; Zhu, M. Generating and Evolving Reward Functions for Highway Driving with Large Language Models. arXiv 2024, arXiv:2406.10540. [Google Scholar]
- Wei, Y.; Shan, X.; Li, J. LERO: LLM-driven Evolutionary framework with Hybrid Rewards and Enhanced Observation for Multi-Agent Reinforcement Learning. arXiv 2025, arXiv:2503.21807. [Google Scholar]
- Jiang, Z.; Schmidt, D.; Srikanth, D.; Xu, D.; Kaplan, I.; Jacenko, D.; Wu, Y. AIDE: AI-Driven Exploration in the Space of Code. arXiv 2025, arXiv:2502.13138. [Google Scholar]
- Pepe, A.; Lin, C.Y.; Magka, D.; Acun, B.; Wu, Y.N.; Protopopov, A.; Wu, C.J.; Bachrach, Y. Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Design, 2026. arXiv arXiv:cs.
- Jeddi, A.; Le, M.N.; Karaimer, H.C.; Derpanis, K.G.; Taati, B. GEAR: Genetic AutoResearch for Agentic Code Evolution. arXiv 2026, arXiv:cs. [Google Scholar]
- Jiang, J.; Zhu, H.; Zhu, Z. SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution, 2026. arXiv arXiv:cs.
- Si, C.; Yang, Z.; Choi, Y.; Candès, E.; Yang, D.; Hashimoto, T. Towards Execution-Grounded Automated AI Research, 2026. arXiv arXiv:cs.
- Yuan, S.; Chen, Z.; Xi, Z.; Ye, J.; Du, Z.; Chen, J. Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training. arXiv 2025, arXiv:cs. [Google Scholar]
- Chen, M.; Lv, C.C.; Zhang, G.; Chang, H.; Zhou, S. HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems. 2026. [Google Scholar]
- Iacob, A.; Jovanovic, A.; Shen, W.F.; Burkhardt, D.; Kurmanji, M.; Tastan, N.; Sani, L.; Venanzi, N.A.E.; Odonnat, A.; Cao, Z.; et al. The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators, 2026. arXiv arXiv:cs.
- Hu, J.; Shukla, P.; Huang, K. Verifying the Verifiers: Failure Attribution for Agentic Benchmark Diagnostics and Training Data Curation. In Proceedings of the ICLR 2026 Workshop on AI with Recursive Self-Improvement (RSI), 2026. [Google Scholar]
- Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models, 2023. arXiv arXiv:cs.
- Chen, X.; Aksitov, R.; Alon, U.; Ren, J.; Xiao, K.; Yin, P.; Prakash, S.; Sutton, C.; Wang, X.; Zhou, D. Universal Self-Consistency for Large Language Model Generation. arXiv 2023, arXiv:cs. [Google Scholar]
- Wang, Z.; Wang, K.; Wang, Q.; Zhang, P.; Li, L.; Yang, Z.; Jin, X.; Yu, K.; Nguyen, M.N.; Liu, L.; et al. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning. arXiv 2025, arXiv:cs. [Google Scholar]
- Yi, B.; Liu, Q.; Cheng, Y.; Xu, H. Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence. arXiv 2026, arXiv:stat. [Google Scholar]
- Xia, P.; Chen, J.; Wang, H.; Liu, J.; Zeng, K.; Wang, Y.; Han, S.; Zhou, Y.; Zhao, X.; Chen, H.; et al. SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning. CoRR 2026, abs/2602.08234, [2602.08234. [Google Scholar] [CrossRef]
- Li, Y.; Miao, R.; Qi, Z.; Lan, T. ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning. CoRR 2026, abs/2603.16060, [2603.16060. [Google Scholar] [CrossRef]
- Hebbar, P.; Manawat, Y.; Verboomen, S.; Ivanova, A.; Palanimalai, S.; Bhatia, K.; Baskaran, V. SIA: Self Improving AI with Harness & Weight Updates. CoRR 2026, abs/2605.27276, [2605.27276. [Google Scholar] [CrossRef]
- Li, Y.; Zhang, Y.; Zhang, X.; Liu, X.; Liu, Y. CODESKILL: Learning Self-Evolving Skills for Coding Agents. CoRR 2026, abs/2605.25430, [2605.25430. [Google Scholar] [CrossRef]
- Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; Anandkumar, A. Voyager: An Open-Ended Embodied Agent with Large Language Models. CoRR 2023, abs/2305.16291, [2305.16291. [Google Scholar] [CrossRef]
- Xia, C.S.; Wang, Z.; Yang, Y.; Wei, Y.; Zhang, L. Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? CoRR 2025, abs/2511.13646, [2511.13646. [Google Scholar] [CrossRef]
- Lou, X.; Lázaro-Gredilla, M.; Dedieu, A.; Wendelken, C.; Lehrach, W.; Murphy, K.P. AutoHarness: improving LLM agents by automatically synthesizing a code harness. abs/2603.03329; CoRR. 2026; p. 2603.03329. [Google Scholar] [CrossRef]
- Fang, T.; Zhang, H.; Zhang, Z.; Ma, K.; Yu, W.; Mi, H.; Yu, D. WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model. arXiv 2025, arXiv:cs. [Google Scholar]
- Guo, Y.; Lee, T.; Shi, L.X.; Chen, J.; Liang, P.; Finn, C. VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model, 2026. arXiv arXiv:cs.








| Level | Category | Programmatic Improvement Loop | Recursion | Domain | ||
| Proposer | Modifier | Validator | ImpModification | Generality | ||
| L1 | Manual Improvement | |||||
| L2 | Assisted | |||||
| Improvement | ||||||
| L3 | Programmatic | |||||
| Self-Improvement | ||||||
| L4 | Bounded Recursive | |||||
| Self-Improvement | ||||||
| L5 | General Recursive | |||||
| Self-Improvement | ||||||
| Work | Self-Improvement Process | Level | ||
| Object | Evidence | Mechanism | ||
| Inner-Loop Trainer Adaptation | ||||
| ACE[270] | Test generator | Code–test execution matrix | Preference optimization | L3 |
| AdaLRS[312] | Learning rate | Loss-descent velocity | Probe with retain/rollback | L3 |
| ADMIRE[279] | Milestone criteria | Successful trajectories | Milestone distillation | L3 |
| ASL[294] | Reward model | Rule-verified outcomes | Alternating role RL | L3 |
| AI Training Manager[311] | Training recipe | Training telemetry | Schema-bounded LLM edits | L3 |
| ARBOR[278] | Rubric memory | Trajectory contrasts | Rubric write–prune | L3 |
| ARCO[285] | Rubric generator | Binary terminal outcomes | Joint rubric–policy RL | L3 |
| ATLAS[310] | DPO coef.; reference anchor | Telemetry; reference KL | Bounded update; anchor swap | L3 |
| CoNL[295] | Critic | Critique-induced gains | Segment-rewarded RL | L3 |
| COPUS[314] | Batch size; parallelism | Gradient noise; throughput | Goodput-driven resharding | L3 |
| CoVerRL[296] | Verifier | Consensus pseudo-labels | Alternating role RL | L3 |
| DR Tulu[276] | Rubric memory | Trajectory contrasts | Rubric write–prune | L3 |
| DYNAMIX[319] | Batch-size controller | Training + system metrics | PPO controller training | L3 |
| EARL[321] | Parallelism config. | Context length; system load | Profile lookup; live switch | L3 |
| ECHO[291] | Critic | Critique-induced gains | Gain-rewarded RL | L3 |
| EvoLM[282] | Rubric generator | Temporal contrasts; margins | Alternating role RL | L3 |
| EPI[313] | Parameter mask | Per-param squared gradients | Periodic mask refresh | L3 |
| EvoRubric[284] | Rubric generator; archive | Rubric validity; consensus | Role RL; rubric archiving | L3 |
| EvoRubrics[281] | Rubric generator | Cross-evaluation scores | Adversarial GRPO update | L3 |
| ICRL[292] | Critic | Critique-induced gains | Alternating role RL | L3 |
| JarvisEvo[290] | Reward model | Human target assessments | Alternating role RL | L3 |
| KungFu[315] | Batch size; workers; topology | Gradient noise; throughput | Rule-based reconfiguration | L3 |
| MagicGUI-RMS[289] | Reward model | Reward-model disagreements | Data write-back; RM retraining | L3 |
| Mutual-Taught[288] | Reward model | Pre/post-update preferences | Alternating DPO updates | L3 |
| ONES[318] | Batch size; GPU allocation | Epoch progress; cluster state | Evolutionary scheduling | L3 |
| Optimus[322] | Resource allocation | Loss history; throughput | Model refit; reallocation | L3 |
| Pollux[316] | Batch size; GPU allocation | Gradient noise; throughput | Goodput co-optimization | L3 |
| Q-Evolve[301] | Value critic | Transitions; terminal reward | IQL critic training | L3 |
| RFT-FM[320] | Training configuration | Telemetry fault fingerprints | Diagnosis-guided repair | L3 |
| RLAC[293] | Critic | Validator outcomes | Alternating DPO updates | L3 |
| RLAnything[300] | Process reward model | Terminal outcomes; consistency | Alternating RM–policy RL | L3 |
| RLAR[286] | Reward-tool library | Tool verification results | Synthesize, verify, commit | L3 |
| RLCER[283] | Rubric generator | Rubric–answer correlation | Correlation-rewarded RL | L3 |
| RL Tango[299] | Verifier | Final-answer correctness | Interleaved GRPO updates | L3 |
| rStar-Math[302] | Process preference model | MCTS values; code execution | Pairwise preference training | L3 |
| Rubick[323] | Execution plan | Measured throughput | Model refit; reconfiguration | L3 |
| RubricEM[277] | Rubric memory; meta-policy | Trajectory contrasts | Write–prune; reflection RL | L3 |
| SAVE[287] | Reward model | On-policy reward–value gaps | Value-anchored self-training | L3 |
| SCORE[297] | Evaluator | Cross-rollout consistency | Consistency-rewarded RL | L3 |
| Sia[317] | Batch size; GPU allocation | Gradient noise; cluster load | Goodput co-optimization | L3 |
| Skill-SD[308] | Teacher state; skill bank | Trajectory rewards; utility | Skill writing; teacher sync | L3 |
| SPARD[306] | Reward weights | Reward progress; dispersion | Mirror-descent reweighting | L3 |
| StepORLM[303] | Process reward model | Solver outcomes; critiques | Iterative RM fine-tuning | L3 |
| STRIDE[304] | Verifier | Final-answer correctness | Verifier fine-tuning | L3 |
| Evaluative Thinking[280] | Evaluation rubric | Responses; RM scores | Meta-reward-model rewrite | L3 |
| UCOB[309] | Teacher-side skill memory | Return gaps; skill utility | Credit-aware skill write–prune | L3 |
| UI-Genie[305] | Reward model | Outcome/continuation labels | Iterative RM fine-tuning | L3 |
| URPO[298] | Reward model | Preference rankings | Rank-rewarded GRPO update | L3 |
| ZeroCoder[307] | Test generator; selector prior | Code–test execution matrix | Role RL; prior recalibration | L3 |
| Outer-Loop Trainer Search through Experiments | ||||
| AIDE[353] | Training script | Validation score; run errors | LLM tree search | L3 |
| AIRA-Design[354] | Training script | Validation loss; executability | LLM tree search | L3 |
| Auto Research[328] | Training script | Eval scores; crashes; runtime | LLM code edits; promotion | L3 |
| AutoTrainess[326] | Fine-tuning recipe | Benchmark scores | LLM recipe assembly | L3 |
| CARD[335] | Reward code | Trajectory preferences | LLM code edits; gated trials | L3 |
| CliffSearch[329] | Optimizer code | Validation loss; reviews | Reviewer-gated evolution | L3 |
| DiscoPOP[348] | Loss function | Post-training scores | LLM generation; selection | L3 |
| LLM-RL HPO[347] | Hyperparameters | Reward/KL dynamics; cost | Multi-fidelity Bayesian opt. | L3 |
| Eureka[334] | Reward code | Fitness; component traces | Evolution; elite retention | L3 |
| FT-Dojo[327] | Fine-tuning recipe | Validation errors; loss curves | LLM proposals; gated trials | L3 |
| Highway Reward Evolution[351] | Reward code | Driving metrics | Evolution with RL trials | L3 |
| LaRes[340] | Reward code | Task success; learning curves | Evolution; elite retention | L3 |
| LERO[352] | Reward code | Team returns; convergence | Evolution with RL trials | L3 |
| LLM-ALSO[336] | Reward-shaping config. | Returns; branch stability | Fork-validated promotion | L3 |
| LLMZero[324] | Training configuration | Validation telemetry | LLM tree search | L3 |
| MLEvolve[338] | Training pipeline | Task metric; run failures | Graph search with memory | L3 |
| OptiCo[330] | Distributed-training config. | Throughput; runtime failures | Multi-agent search; memory | L3 |
| POISE[333] | Policy-optimization code | Dynamics; held-out scores | Archive-guided LLM edits | L3 |
| R*[341] | Reward code; coefficients | Fitness; component traces | Evolution with RL trials | L3 |
| REvolve[337] | Reward code | Human rollout preferences | Island evolution; migration | L3 |
| RF-Agent[342] | Reward code | Return; component feedback | LLM tree search | L3 |
| RHyVE[346] | Reward schedule | Fork margins; agreement | Fork-validated promotion | L3 |
| ROSKA[343] | Reward code; fusion ratio | Fitness; component traces | Evolution; tuned policy fusion | L3 |
| Reward Synthesis[350] | Reward code | Executability; task F1 | Ranked generation; ensemble | L3 |
| Self-Evolving Recommender[339] | Training + reward code | Offline loss; A/B metrics | LLM code edits; staged gates | L3 |
| SMCEvolve[356] | Training script | Validation loss; run failures | SMC resampling; mutation | L3 |
| TREX[332] | Fine-tuning recipe | Validation score; run failures | LLM tree search | L3 |
| Meta-Loop Improvement-Mechanism Evolution | ||||
| AutoScientists[331] | Role-allocation policy | Validation deltas; stagnation | Evidence-triggered reorg. | L4 |
| Bilevel Autoresearch[238] | Search runner code | Proposal–outcome traces | LLM-generated replacement | L4 |
| DERL[349] | Reward meta-optimizer | Inner-policy validation | GRPO meta-training | L4 |
| EvoTrainer[325] | Diagnostic harness | Diagnostic gaps; branch scores | Retained harness revision | L4 |
| GEAR[355] | Search controller code | Crossover failures | Source-level repair | L4 |
| Execution-Grounded AI Research[357] | Ideation policy | Validation score; run status | Execution-rewarded RL | L4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).